← Search

Stefanos Nikolaidis

38 accepted papers

2026

AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity Optimization

ICLR 2026poster

Quality-Diversity (QD) algorithms have shown remarkable success in discovering diverse, high-performing solutions, but rely heavily on hand-crafted behavioral descriptors that constrain exploration to predefined notions of diversity. Leveraging the equivalence between policies and occupancy measures…

Cited by 0SourceScholar
2026

Discount Model Search for Quality Diversity Optimization in High-Dimensional Measure Spaces

ICLR 2026oral

Quality diversity (QD) optimization searches for a collection of solutions that optimize an objective while attaining diverse outputs of a user-specified, vector-valued measure function. Contemporary QD algorithms are typically limited to low-dimensional measures because high-dimensional measures ar…

Cited by 0SourcecodeScholar
2026

Fix the Mind, Not the Move: Interpretable AI Assistance via Knowledge-Gap Localization

ICML 2026poster

AI assistants in human-AI collaboration often correct suboptimal human actions through behavioral feedback (e.g., alerts or steering-wheel nudges in assistive driving). Such interventions can mitigate immediate errors, but long-term improvement requires addressing the underlying misconceptions that …

Cited by 0SourceScholar
2025

Adaptively Coordinating with Novel Partners via Learned Latent Strategies

NeurIPS 2025poster

Adaptation is the cornerstone of effective collaboration among heterogeneous team members. In human-agent teams, artificial agents need to adapt to their human partners in real time, as individuals often have unique preferences and policies that may change dynamically throughout interactions. This b…

Cited by 0SourceScholar
2025

Integrating Field of View in Human-Aware Collaborative Planning

ICRA 2025

In human-robot collaboration (HRC), it is crucial for robot agents to consider humans' knowledge of their surroundings. In reality, humans possess a narrow field of view (FOV), limiting their perception. However, research on HRC often overlooks this aspect and presumes an omniscient human collaborat

Cited by 3SourceScholar
2025

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

CoRL 2025poster

Vision-Language Models (VLMs) have revolutionized artificial intelligence and robotics due to their commonsense reasoning capabilities. In robotic manipulation, VLMs are used primarily as high-level planners, but recent work has also studied their lower-level reasoning ability, which refers to makin…

Cited by 0SourceScholar
2024

BayRnTune: Adaptive Bayesian Domain Randomization via Strategic Fine-tuning

IROS 2024poster

Domain randomization (DR), which entails training a policy with randomized dynamics, has proven to be a simple yet effective algorithm for reducing the gap between simulation and the real world. However, DR often requires careful tuning of randomization parameters. Methods like Bayesian Domain Rando…

Cited by 3SourceScholar
2024

Enabling Adaptive Agent Training in Open-Ended Simulators by Targeting Diversity

NeurIPS 2024poster

The wider application of end-to-end learning methods to embodied decision-making domains remains bottlenecked by their reliance on a superabundance of training data representative of the target domain. Meta-reinforcement learning (meta-RL) approaches abandon the aim of zero-shot *generalization*—the…

2024

Guidance Graph Optimization for Lifelong Multi-Agent Path Finding

IJCAI 2024poster

We study how to use guidance to improve the throughput of lifelong Multi-Agent Path Finding (MAPF). Previous studies have demonstrated that, while incorporating guidance, such as highways, can accelerate MAPF algorithms, this often results in a trade-off with solution quality. In addition, how to ge…

2024

Multi-Robot Task Allocation Under Uncertainty Via Hindsight Optimization

ICRA 2024poster

Multi-robot systems are becoming increasingly prevalent in various real-world applications, such as manufacturing and warehouse logistics. These systems face complex challenges in 1) task allocation due to factors like time-extended tasks, and agent specialization, and 2) uncertainties in task execu…

Cited by 2SourceScholar
2024

Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning

ICLR 2024spotlight

Training generally capable agents that thoroughly explore their environment and learn new and diverse skills is a long-term goal of robot learning. Quality Diversity Reinforcement Learning (QD-RL) is an emerging research area that blends the best aspects of both fields – Quality Diversity (QD) provi…

Cited by 15SourcePDFScholar
2024

Quality-Diversity Generative Sampling for Learning with Synthetic Data

AAAI 2024technical

Generative models can serve as surrogates for some real data sources by creating synthetic training datasets, but in doing so they may transfer biases to downstream tasks. We focus on protecting quality and diversity when generating synthetic training datasets. We propose quality-diversity generativ…

2024

Selecting Source Tasks for Transfer Learning of Human Preferences

RA-L 2024

We address the challenge of transferring human preferences for action selection from simpler source tasks to complex target tasks. Our goal is to enable robots to support humans proactively by predicting their actions — without requiring demonstrations of their preferred action sequences in the targ

Cited by 0SourceScholar
2023

Arbitrarily Scalable Environment Generators via Neural Cellular Automata

NeurIPS 2023poster

We study the problem of generating arbitrarily large environments to improve the throughput of multi-robot systems. Prior work proposes Quality Diversity (QD) algorithms as an effective method for optimizing the environments of automated warehouses. However, these approaches optimize only relatively…

2023

Contingency-Aware Task Assignment and Scheduling for Human-Robot Teams

ICRA 2023poster

We consider the problem of task assignment and scheduling for human-robot teams to enable the efficient completion of complex problems, such as satellite assembly. In high-mix, low volume settings, we must enable the human-robot team to handle uncertainty due to changing task requirements, potential…

Cited by 11SourceScholar
2023

Inverse Reinforcement Learning Framework for Transferring Task Sequencing Policies from Humans to Robots in Manufacturing Applications

ICRA 2023poster

In this work, we present an inverse reinforcement learning approach for solving the problem of task sequencing for robots in complex manufacturing processes. Our proposed framework is adaptable to variations in process and can perform sequencing for entirely new parts. We prescribe an approach to ca…

Cited by 15SourceScholar
2023

Learning Performance Graphs From Demonstrations via Task-Based Evaluations

RA-L 2023

In the paradigm of robot learning-from-demonstra tions (LfD), understanding and evaluating the demonstrated behaviors plays a critical role in extracting control policies for robots. Without this knowledge, a robot may infer incorrect reward functions that lead to undesirable or unsafe control polic

Cited by 5SourceScholar
2023

Multi-Robot Coordination and Layout Design for Automated Warehousing

IJCAI 2023poster

With the rapid progress in Multi-Agent Path Finding (MAPF), researchers have studied how MAPF algorithms can be deployed to coordinate hundreds of robots in large automated warehouses. While most works try to improve the throughput of such warehouses by developing better MAPF algorithms, we focus on…

2023

PATO: Policy Assisted TeleOperation for Scalable Robot Data Collection

RSS 2023poster

Large-scale data is an essential component of machine learning as demonstrated in recent advances in natural language processing and computer vision research. However, collecting large-scale robotic data is much more expensive and slower as each operator can control only a single robot at a time. To…

Cited by 19SourcePDFScholar
2023

Surrogate Assisted Generation of Human-Robot Interaction Scenarios

CoRL 2023oral

As human-robot interaction (HRI) systems advance, so does the difficulty of evaluating and understanding the strengths and limitations of these systems in different environments and with different users. To this end, previous methods have algorithmically generated diverse scenarios that reveal syste…

Cited by 11SourcecodeScholar
2023

Training Diverse High-Dimensional Controllers by Scaling Covariance Matrix Adaptation MAP-Annealing

RA-L 2023

Pre-training a diverse set of neural network controllers in simulation has enabled robots to adapt online to damage in robot locomotion tasks. However, finding diverse, high-performing controllers requires expensive network training and extensive tuning of a large number of hyperparameters. On the o

Cited by 16SourcecodeScholar
2022

Deep Surrogate Assisted Generation of Environments

NeurIPS 2022accept

Recent progress in reinforcement learning (RL) has started producing generally capable agents that can solve a distribution of complex environments. These agents are typically tested on fixed, human-authored environments. On the other hand, quality diversity (QD) optimization has been proven to be a…

2021

A Quality Diversity Approach to Automatically Generating Human-Robot Interaction Scenarios in Shared Autonomy

RSS 2021poster

The growth of scale and complexity of interactions between humans and robots highlights the need for new computational methods to automatically evaluate novel algorithms and applications. Exploring diverse scenarios of humans and robots interacting in simulation can improve understanding of the robo…

Cited by 50SourcePDFScholar
2021

Autonomy in Physical Human-Robot Interaction: A Brief Survey

RA-L 2021

Sharing the control of a robotic system with an autonomous controller allows a human to reduce his/her cognitive and physical workload during the execution of a task. In recent years, the development of inference and learning techniques has widened the spectrum of applications of shared control (SC)

Cited by 193SourceScholar
2021

Design and Evaluation of a Hair Combing System Using a General-Purpose Robotic Arm

IROS 2021poster

This work introduces an approach for automatic hair combing by a lightweight robot. For people living with limited mobility, dexterity, or chronic fatigue, combing hair is often a difficult task that negatively impacts personal routines. We propose a modular system for enabling general robot manipul…

Cited by 8SourceScholar
2021

Illuminating Mario Scenes in the Latent Space of a Generative Adversarial Network

AAAI 2021technical

Generative adversarial networks (GANs) are quickly becoming a ubiquitous approach to procedurally generating video game levels. While GAN generated levels are stylistically similar to human-authored examples, human designers often want to explore the generative design space of GANs to extract intere…

2021

Learning Collaborative Pushing and Grasping Policies in Dense Clutter

ICRA 2021poster

Robots must reason about pushing and grasping in order to engage in flexible manipulation in cluttered environments. Earlier works on learning pushing and grasping only consider each operation in isolation or are limited to top-down grasping and bin-picking. We train a robot to learn joint planar pu…

Cited by 40SourceScholar
2021

Learning From Demonstrations Using Signal Temporal Logic in Stochastic and Continuous Domains

RA-L 2021

Learning control policies that are safe, robust and interpretable are prominent challenges in developing robotic systems. Learning-from-demonstrations with formal logic is an arising paradigm in reinforcement learning to estimate rewards and extract robot control policies that seek to overcome these

Cited by 33SourceScholar
2021

Robotic Lime Picking by Considering Leaves as Permeable Obstacles

IROS 2021poster

The problem of robotic lime picking is challenging; lime plants have dense foliage which makes it difficult for a robotic arm to grasp a lime without coming in contact with leaves. Existing approaches either do not consider leaves, or treat them as obstacles and completely avoid them, often resultin…

Cited by 19SourceScholar
2021

Two-Stage Clustering of Human Preferences for Action Prediction in Assembly Tasks

ICRA 2021poster

To effectively assist human workers in assembly tasks a robot must proactively offer support by inferring their preferences in sequencing the task actions. Previous work has focused on learning the dominant preferences of human workers for simple tasks largely based on their intended goal. However,…

Cited by 13SourceScholar
2020

Fair Contextual Multi-Armed Bandits: Theory and Experiments

UAI 2020poster

When an AI system interacts with multiple users, it frequently needs to make allocation decisions. For instance, a virtual agent decides whom to pay attention to in a group, or a factory robot selects a worker to deliver a part.Demonstrating fairness in decision making is essential for such systems…

Cited by 79SourcePDFScholar
2019

Robot Object Referencing through Legible Situated Projections

ICRA 2019poster

The ability to reference objects in the environment is a key communication skill that robots need for complex, task-oriented human-robot collaborations. In this paper we explore the use of projections, which are a powerful communication channel for robot-to-human information transfer as they allow f…

Cited by 19SourceScholar