Back to advisors
KO
About
Despite many large breakthroughs in Artificial Intelligence in the last decade, robots are still struggling to solve tasks that we as humans take for granted, like emptying a dishwasher or learning multiple skills in a sequence. In this talk, Kai Olav Ellefsen will talk about work he and colleagues at the group of Robotics and Intelligent Systems (ROBIN), University of Oslo, have done with the goal of making robots more robust and better learners, by taking inspiration from how humans and animals learn.
Selected publications
- Scientific articles and book chapters
- Kvalsund, Mia-Katrin Ose; Ellefsen, Kai Olav; Glette, Kyrre; Pontes-Filho, Sidney & Lepperød, Mikkel Elle (2026). Sensor movement drives emergent attention and scalability in active neural cellular automata. Neural Networks. ISSN 0893-6080. 200. doi: 10.1016/j.neunet.2026.108798. Full text in Research Archive
- Lømo, Tobias; Baselizadeh, Adel; Ellefsen, Kai Olav & Tørresen, Jim (2026). Dual Process Dreamer: Fast and Slow Decision-Making with World Models. Proceedings of the International Conference on Agents and Artificial Intelligence (ICAART). ISSN 2184-3589. 2, p. 1230–1241. doi: 10.5220/0014243200004052. Full text in Research Archive Show summary Most robot systems are based on a single decision-making process. This process needs to balance time, energy, and accuracy in every situation. However, according to ”dual process theory” (DPT) from cognitive psychology, this is not how humans work. Depending on the situation, we have the ability to switch between two thinking methods, a fast system 1 (S1) and a slower system 2 (S2). In this paper, we propose a novel approach to a dual process architecture for robots and agents. Our method, called Dual Process Dreamer (DPDreamer), is a combination of a reinforcement learning policy network, a planning algorithm, and a learned world model. The world model allows the parts of DPDreamer to work together and create a more integrated system compared to previous proposals of DPT systems. DPDreamer was tested in a puzzle game called Sokoban, and by balancing the use of S1 and S2, DPDreamer managed a success rate similar to S2 while using S1 most of the time, showing the benefit of using a more adaptable system.
- Watanabe, Shin; Horn, Geir; Tørresen, Jim & Ellefsen, Kai Olav (2025). Integrating Bilevel Planning and Offline Skill Learning for Enhancing Mobile Manipulation, 2025 IEEE 21st International Conference on Automation Science and Engineering (CASE). Institute of Electrical and Electronics Engineers (IEEE). ISSN 9798331522476. p. 2275–2280. doi: 10.1109/CASE58245.2025.11163881. Full text in Research Archive Show summary Solving complex robotic mobile manipulation tasks requires both planning the sequence of skills to execute and learning how to robustly execute each skill. Planning-based approaches such as task and motion planning (TAMP) can help train skills more efficiently through demonstrations, while learning-based approaches such as reinforcement learning (RL) can help plan tasks more quickly through heuristics. This paper presents a novel approach to generalizing mobile manipulation tasks by synergistically combining sampling-based TAMP and value-based RL. The TAMP solver first generates suboptimal demonstration trajectories of a particular skill, from which an offline RL algorithm distills a robust policy and a value function, the latter serving as a skill feasibility classifier. The policy and the classifier are both fed back into the TAMP workflow not only to improve the skill success rate but also to speed up the planner by sampling robot configurations for which the skill is likely to succeed. We evaluate the approach on a simulated block-pushing domain. Re-purposing a byproduct of an offline skill-learning process leads to an integrated planning and learning system that exploits the awareness of its own skill competence.
- Bruin, Ege de; Glette, Kyrre & Ellefsen, Kai Olav (2025). Integrating Sample Inheritance into Bayesian Optimization for Evolutionary Robotics. In Witkowski, Olaf; Adams, Alyssa M.; Sinapayen, Lana; Baltieri, Manuel & Khosravy, Mahdi (Ed.), ALIFE 2025: Ciphers of Life: Proceedings of the Artificial Life Conference 2025. MIT Press. doi: 10.1162/ISAL.a.866. Full text in Research Archive Show summary In evolutionary robotics, robot morphologies are designed automatically using evolutionary algorithms. This creates a body-brain optimization problem, where both morphology and control must be optimized together. A common approach is to include controller optimization for each morphology, but starting from scratch for every new body may require a high controller learning budget. We address this by using Bayesian optimization for controller optimization, exploiting its sample efficiency and strong exploration capabilities, and using sample inheritance as a form of Lamarckian inheritance. Under a deliberately low controller learning budget for each morphology, we investigate two types of sample inheritance: (1) transferring all the parent’s samples to the offspring to be used as prior without evaluating them, and (2) reevaluating the parent’s best samples on the offspring. Both are compared to a baseline without inheritance. Our results show that reevaluation performs best, with prior-based inheritance also outperforming no inheritance. Analysis reveals that while the learning budget is too low for a single morphology, generational inheritance compensates for this by accumulating learned adaptations across generations. Furthermore, inheritance mainly benefits offspring morphologies that are similar to their parents. Finally, we demonstrate the critical role of the environment, with more challenging environments resulting in more stable walking gaits. Our findings highlight that inheritance mechanisms can boost performance in evolutionary robotics without needing large learning budgets, offering an efficient path toward more capable robot design.
- Nergård, Katrine Linnea; Ellefsen, Kai Olav & Tørresen, Jim (2025). Fast or Slow: Adaptive Decision Making in Reinforcement Learning with Pre-Trained LLMs. In Ugur, Emre; Sciutti, Alessandra & Rohlfing, Katharina (Ed.), 2025 IEEE International Conference on Development and Learning (ICDL). IEEE (Institute of Electrical and Electronics Engineers). ISSN 9798331543433. doi: 10.1109/ICDL63968.2025.11204357. Full text in Research Archive
- Waarum, Ivar-Kristian; Blomberg, Ann Elisabeth Albright; Krogstad, Thomas Røbekk; Dewar, Marius & Ellefsen, Kai Olav (2025). Optimal Sampling Patterns for Robotic Environmental Monitoring. International Conference on Control, Automation and Robotics (ICCAR). ISSN 2251-2446. 2025(2025), p. 163–169. doi: 10.1109/iccar64901.2025.11073028. Full text in Research Archive Show summary Cost-efficient methods and technologies for environmental monitoring is an enabler for holistic decision making and sustainable management of natural resources. The information value from monitoring surveys with robotic sensor platforms such as AUVs and UAVs can increase if domain knowledge is exploited when selecting the sampling locations. We investigate how to maximize the information value of surveys targeting dispersible substances by adjusting the survey pattern in accordance with the direction and strength of the wind or water currents. Samples are acquired with a range of different patterns, deployed in a large number of randomly generated emission scenarios in a Monte Carlo framework. The information value from each survey is used as a performance metric to identify opportune pattern configurations. Among the results is the impact that the orientation of the survey pattern relative to the direction of the wind has on the information value, and that given proper orientation the line spacing can be wider. The results are used to suggest guidelines for robotic environmental monitoring surveys.
- Bruin, Ege de; Glette, Kyrre & Ellefsen, Kai Olav (2025). Generational Replacement and Learning for High-Performing and Diverse Populations in Evolvable Robots. In Obafemi-Ajayi, Tayo (Eds.), 2025 IEEE Symposium on Computational Intelligence in Artificial Life and Cooperative Intelligent Systems. IEEE (Institute of Electrical and Electronics Engineers). ISSN 9798331508388. doi: 10.1109/alife-cis64968.2025.10979828. Full text in Research Archive Show summary Evolutionary Robotics offers the possibility to design robots to solve a specific task automatically by optimizing their morphology and control together. However, this co-optimization of body and control is challenging, because controllers need some time to adapt to the evolving morphology - which may make it difficult for new and promising designs to enter the evolving population. A solution to this is to add intra-life learning, defined as an additional controller optimization loop, to each individual in the evolving population. A related problem is the lack of diversity often seen in evolving populations as evolution narrows the search down to a few promising designs too quickly. This problem can be mitigated by implementing full generational replacement, where offspring robots replace the whole population. This solution for increasing diversity usually comes at the cost of lower performance compared to using elitism. In this work, we show that combining such generational replacement with intra-life learning can increase diversity while retaining performance. We also highlight the importance of performance metrics when studying learning in morphologically evolving robots, showing that evaluating according to function evaluations versus according to generations of evolution can give different conclusions.
Data verified 9/6/2026Source