← Search

Pavel Kolev

8 accepted papers

2025

SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models

ICML 2025poster

Exploration is a cornerstone of reinforcement learning (RL). Intrinsic motivation attempts to decouple exploration from external, task-based rewards. However, established approaches to intrinsic motivation that follow general principles such as information gain, often only uncover low-level interact…

Cited by 2SourcePDFScholar
2024

Learning Diverse Skills for Local Navigation under Multi-constraint Optimality

ICRA 2024poster

Despite many successful applications of data-driven control in robotics, extracting meaningful diverse behaviors remains a challenge. Typically, task performance needs to be compromised in order to achieve diversity. In many scenarios, task requirements are specified as a multitude of reward terms,…

Cited by 7SourceScholar
2023

Benchmarking Offline Reinforcement Learning on Real-Robot Hardware

ICLR 2023top-25%

Learning policies from previously recorded data is a promising direction for real-world robotics tasks, as online learning is often infeasible. Dexterous manipulation in particular remains an open problem in its general form. The combination of offline reinforcement learning with large diverse datas…

Cited by 38SourcePDFScholar
2023

On Imitation in Mean-field Games

NeurIPS 2023poster

We explore the problem of imitation learning (IL) in the context of mean-field games (MFGs), where the goal is to imitate the behavior of a population of agents following a Nash equilibrium policy according to some unknown payoff function. IL in MFGs presents new challenges compared to single-agent…

Cited by 2SourcePDFScholar
2023

Versatile Skill Control via Self-supervised Adversarial Imitation of Unlabeled Mixed Motions

ICRA 2023poster

Learning diverse skills is one of the main challenges in robotics. To this end, imitation learning approaches have achieved impressive results. These methods require explicitly labeled datasets or assume consistent skill execution to enable learning and active control of individual behaviors, which…

Cited by 34SourceScholar
2020

Secretary and Online Matching Problems with Machine Learned Advice

NeurIPS 2020poster

The classical analysis of online algorithms, due to its worst-case nature, can be quite pessimistic when the input instance at hand is far from worst-case. Often this is not an issue with machine learning approaches, which shine in exploiting patterns in past inputs in order to predict the future. H…

Cited by 136SourcePDFScholar