← Search

Mykel J Kochenderfer

57 accepted papers

2026

World Model Failure Classification and Anomaly Detection for Autonomous Inspection

ICRA 2026poster

Autonomous inspection robots for monitoring industrial sites can reduce costs and risks associated with human-led inspection. However, accurate readings can be challenging due to occlusions, limited viewpoints, or unexpected environmental conditions. We propose a hybrid framework that combines super…

2025

Enhanced Importance Sampling Through Latent Space Exploration in Normalizing Flows

AAAI 2025technical

Importance sampling is a rare event simulation technique used in Monte Carlo simulations to bias the sampling distribution towards the rare event of interest. By assigning appropriate weights to sampled points, importance sampling allows for more efficient estimation of rare events or tails of distr…

2025

Importance Sampling-Guided Meta-Training for Intelligent Agents in Highly Interactive Environments

RA-L 2025

Training intelligent agents to navigate highly interactive environments presents significant challenges. While guided meta reinforcement learning (RL) approach that first trains a guiding policy to train the ego agent has proven effective in improving generalizability across scenarios with various l

Cited by 4SourceScholar
2025

SayComply: Grounding Field Robotic Tasks in Operational Compliance Through Retrieval-Based Language Models

ICRA 2025

This paper addresses the problem of task planning for robots that must comply with operational manuals in real-world settings. Task planning under these constraints is essential for enabling autonomous robot operation in domains that require adherence to domain-specific knowledge. Current methods fo

Cited by 6SourcecodeScholar
2025

Semi-Markovian Planning to Coordinate Aerial and Maritime Medical Evacuation Platforms

AAAI 2025technical

The transfer of patients between two aircraft using an underway watercraft increases medical evacuation reach and flexibility in maritime environments. The selection of any one of multiple underway watercraft for patient exchange is complicated by participating aircraft utilization histories and par…

Cited by 0SourcePDFScholar
2024

Constrained Hierarchical Monte Carlo Belief-State Planning

ICRA 2024poster

Optimal plans in Constrained Partially Observable Markov Decision Processes (CPOMDPs) maximize reward objectives while satisfying hard cost constraints, generalizing safe planning under state and transition uncertainty. Unfortunately, online CPOMDP planning is extremely difficult in large or continu…

Cited by 3SourcecodeScholar
2024

ConstrainedZero: Chance-Constrained POMDP Planning Using Learned Probabilistic Failure Surrogates and Adaptive Safety Constraints

IJCAI 2024poster

To plan safely in uncertain environments, agents must balance utility with safety constraints. Safe planning problems can be modeled as a chance-constrained partially observable Markov decision process (CC-POMDP) and solutions often use expensive rollouts or heuristics to estimate the optimal value…

2024

Disentangled Neural Relational Inference for Interpretable Motion Prediction

RA-L 2024

Effective interaction modeling and behavior prediction of dynamic agents play a significant role in interactive motion planning for autonomous robots. Although existing methods have improved prediction accuracy, few research efforts have been devoted to enhancing prediction model interpretability an

Cited by 9SourceScholar
2024

Distributed Online Planning for Min-Max Problems in Networked Markov Games

RA-L 2024

Min-max problems are important in multi-agent sequential decision-making because they improve the performance of the worst-performing agent in the network. However, solving the multi-agent min-max problem is challenging. We propose a modular, distributed, online planning-based algorithm that is able

Cited by 1SourcecodeScholar
2024

Scene Informer: Anchor-based Occlusion Inference and Trajectory Prediction in Partially Observable Environments

ICRA 2024poster

Navigating complex and dynamic environments requires autonomous vehicles (AVs) to reason about both visible and occluded regions. This involves predicting the future motion of observed agents, inferring occluded ones, and modeling their interactions based on vectorized scene representations of the p…

Cited by 10SourcecodeScholar
2024

Semantic Belief Behavior Graph: Enabling Autonomous Robot Inspection in Unknown Environments

IROS 2024poster

This paper addresses the problem of autonomous robotic inspection in complex and unknown environments. This capability is crucial for efficient and precise inspections in various real-world scenarios, even when faced with perceptual uncertainty and lack of prior knowledge of the environment. Existin…

Cited by 5SourceScholar
2023

Fast and Scalable Signal Inference for Active Robotic Source Seeking

ICRA 2023poster

In active source seeking, a robot takes repeated measurements in order to locate a signal source in a cluttered and unknown environment. A key component of an active source seeking robot planner is a model that can produce estimates of the signal at unknown locations with uncertainty quantification.…

Cited by 9SourceScholar
2023

Model Predictive Optimized Path Integral Strategies

ICRA 2023poster

We generalize the derivation of model predictive path integral control (MPPI) to allow for a single joint distribution across controls in the control sequence. This reformation allows for the implementation of adaptive importance sampling (AIS) algorithms into the original importance sampling step w…

Cited by 24SourcecodeScholar
2023

SHAIL: Safety-Aware Hierarchical Adversarial Imitation Learning for Autonomous Driving in Urban Environments

ICRA 2023poster

Designing a safe and human-like decision-making system for an autonomous vehicle is a challenging task. Generative imitation learning is one possible approach for automating policy-building by leveraging both real-world and simulated decisions. Previous work that applies generative imitation learnin…

Cited by 22SourcecodeScholar
2023

Safe and Efficient Navigation in Extreme Environments using Semantic Belief Graphs

ICRA 2023poster

To achieve autonomy in unknown and unstruc-tured environments, we propose a method for semantic-based planning under perceptual uncertainty. This capability is cru-cial for safe and efficient robot navigation in environment with mobility-stressing elements that require terrain-specific locomotion po…

Cited by 8SourceScholar
2023

Sequential Bayesian Optimization for Adaptive Informative Path Planning with Multimodal Sensing

ICRA 2023poster

Adaptive Informative Path Planning with Multi-modal Sensing (AIPPMS) considers the problem of an agent equipped with multiple sensors, each with different sensing accuracy and energy costs. The agent's goal is to explore the environment and gather information subject to its resource constraints in u…

Cited by 18SourcecodeScholar
2022

Adaptive Coverage Path Planning for Efficient Exploration of Unknown Environments

IROS 2022poster

We present a method for solving the coverage problem with the objective of autonomously exploring an unknown environment under mission time constraints. Here, the robot is tasked with planning a path over a horizon such that the accumulated area swept out by its sensor footprint is maximized. Becaus…

Cited by 15SourceScholar
2022

Capability-Aware Task Allocation and Team Formation Analysis for Cooperative Exploration of Complex Environments

IROS 2022poster

To achieve autonomy in complex real-world exploration missions, we consider deployment strategies for a team of robots with heterogeneous capabilities. We formulate a multi-robot exploration mission and compute an operation policy to maintain robot team productivity and maximize mission success. The…

Cited by 3SourceScholar
2022

Dynamics-Aware Spatiotemporal Occupancy Prediction in Urban Environments

IROS 2022poster

Detection and segmentation of moving obstacles, along with prediction of the future occupancy states of the local environment, are essential for autonomous vehicles to proactively make safe and informed decisions. In this paper, we propose a framework that integrates the two capabilities together us…

Cited by 18SourceScholar
2022

FIG-OP: Exploring Large-Scale Unknown Environments on a Fixed Time Budget

IROS 2022poster

We present a method for autonomous exploration of large-scale unknown environments under mission time con-straints. We start by proposing the Frontloaded Information Gain Orienteering Problem (FIG-OP) - a generalization of the traditional orienteering problem where the assumption of a reliable envir…

Cited by 22SourceScholar
2022

How Do We Fail? Stress Testing Perception in Autonomous Vehicles

IROS 2022poster

Autonomous vehicles (AVs) rely on environment perception and behavior prediction to reason about agents in their surroundings. These perception systems must be robust to adverse weather such as rain, fog, and snow. However, validation of these systems is challenging due to their complexity and depen…

Cited by 21SourcecodeScholar
2022

Infrastructure-Enabled Autonomy: An Attention Mechanism for Occlusion Handling

ICRA 2022poster

Although there has been tremendous progress in autonomous driving, navigating environments and predicting the behavior of other drivers in the presence of occlusions remains challenging. Cities have started investing in infrastructure sensors that could provide information about occluded spaces. We…

Cited by 6SourceScholar
2022

Learning Emergent Discrete Message Communication for Cooperative Reinforcement Learning

ICRA 2022poster

Communication is an important factor that en-ables agents to work cooperatively in multi-agent reinforcement learning (MARL) contexts. Prior work used continuous message communication whose high representational capacity comes at the expense of interpretability. Allowing agents to learn their own di…

Cited by 19SourceScholar
2022

Multi-Agent Variational Occlusion Inference Using People as Sensors

ICRA 2022poster

Autonomous vehicles must reason about spatial occlusions in urban environments to ensure safety without being overly cautious. Prior work explored occlusion inference from observed social behaviors of road agents, hence treating people as sensors. Inferring occupancy from agent behaviors is an inher…

Cited by 27SourcecodeScholar
2022

Multi-Objective Policy Gradients with Topological Constraints

IROS 2022poster

Multi-objective optimization models that encode ordered sequential constraints provide a solution to model various challenging problems including encoding preferences, modeling a curriculum, and enforcing measures of safety. A recently developed theory of topological Markov decision processes (TMDPs…

Cited by 3SourceScholar
2022

Recursive Reasoning Graph for Multi-Agent Reinforcement Learning

AAAI 2022technical

Multi-agent reinforcement learning (MARL) provides an efficient way for simultaneously learning policies for multiple agents interacting with each other. However, in scenarios requiring complex interactions, existing algorithms can suffer from an inability to accurately anticipate the influence of s…

Cited by 10SourcePDFScholar
2022

Scalable Anytime Planning for Multi-Agent MDPs (Extended Abstract)

IJCAI 2022poster

We present a scalable planning algorithm for multi-agent sequential decision problems that require dynamic collaboration. Teams of agents need to coordinate decisions in many domains, but naive approaches fail due to the exponential growth of the joint action space with the number of agents. We c…

Cited by 0SourcePDFScholar
2021

3D Radar Velocity Maps for Uncertain Dynamic Environments

IROS 2021poster

Future urban transportation concepts include a mixture of ground and air vehicles with varying degrees of autonomy in a congested environment. In such dynamic environments, occupancy maps alone are not sufficient for safe path planning. Safe and efficient transportation requires reasoning about the…

Cited by 5SourcecodeScholar
2021

Bayesian Optimized Monte Carlo Planning

AAAI 2021technical

Online solvers for partially observable Markov decision processes have difficulty scaling to problems with large action spaces. Monte Carlo tree search with progressive widening attempts to improve scaling by sampling from the action space to construct a policy search tree. The performance of progre…

2021

Double-Prong ConvLSTM for Spatiotemporal Occupancy Prediction in Dynamic Environments

ICRA 2021poster

Predicting the future occupancy state of an environment is important to enable informed decisions for autonomous vehicles. Common challenges in occupancy prediction include vanishing dynamic objects and blurred predictions, especially for long prediction horizons. In this work, we propose a double-p…

Cited by 26SourcecodeScholar
2021

Finding Failures in High-Fidelity Simulation using Adaptive Stress Testing and the Backward Algorithm

IROS 2021poster

Validating the safety of autonomous systems generally requires the use of high-fidelity simulators that adequately capture the variability of real-world scenarios. However, it is generally not feasible to exhaustively search the space of simulation scenarios for failures. Adaptive stress testing (AS…

Cited by 30SourcecodeScholar
2021

Improved POMDP Tree Search Planning with Prioritized Action Branching

AAAI 2021technical

Online solvers for partially observable Markov decision processes have difficulty scaling to problems with large action spaces. This paper proposes a method called PA-POMCPOW to sample a subset of the action space that provides varying mixtures of exploitation and exploration for inclusion in a sear…

2021

Reinforcement Learning for Autonomous Driving with Latent State Inference and Spatial-Temporal Relationships

ICRA 2021poster

Deep reinforcement learning (DRL) provides a promising way for learning navigation in complex autonomous driving scenarios. However, identifying the subtle cues that can indicate drastically different outcomes remains an open problem with designing autonomous systems that operate in human environmen…

Cited by 80SourceScholar
2020

Efficient Large-Scale Multi-Drone Delivery Using Transit Networks

ICRA 2020poster

We consider the problem of controlling a large fleet of drones to deliver packages simultaneously across broad urban areas. To conserve energy, drones hop between public transit vehicles (e.g., buses and trams). We design a comprehensive algorithmic framework that strives to minimize the maximum tim…

Cited by 147SourcecodeScholar
2020

Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders

NeurIPS 2020poster

Discrete latent spaces in variational autoencoders have been shown to effectively capture the data distribution for many real-world problems such as natural language understanding, human intent prediction, and visual scene representation. However, discrete latent spaces need to be sufficiently large…

2020

Handling Missing Data with Graph Representation Learning

NeurIPS 2020poster

Machine learning with missing data has been approached in many different ways, including feature imputation where missing feature values are estimated based on observed values and label prediction where downstream labels are learned directly from incomplete data. However, existing imputation models…

2020

Optimal Sequential Task Assignment and Path Finding for Multi-Agent Robotic Assembly Planning

ICRA 2020poster

We study the problem of sequential task assignment and collision-free routing for large teams of robots in applications with inter-task precedence constraints (e.g., task A and task B must both be completed before task C may begin). Such problems commonly occur in assembly planning for robotic manuf…

Cited by 54SourceScholar
2020

Provably Efficient Reward-Agnostic Navigation with Linear Value Iteration

NeurIPS 2020poster

There has been growing progress on theoretical analyses for provably efficient learning in MDPs with linear function approximation, but much of the existing work has made strong assumptions to enable exploration by conventional exploration frameworks. Typically these assumptions are stronger than wh…

Cited by 71SourcePDFScholar
2019

Almost Horizon-Free Structure-Aware Best Policy Identification with a Generative Model

NeurIPS 2019poster

This paper focuses on the problem of computing an $\epsilon$-optimal policy in a discounted Markov Decision Process (MDP) provided that we can access the reward and transition function through a generative model. We propose an algorithm that is initially agnostic to the MDP but that can leverage the…

Cited by 47SourcePDFScholar
2019

EnsembleDAgger: A Bayesian Approach to Safe Imitation Learning

IROS 2019poster

Although imitation learning is often used in robotics, the approach frequently suffers from data mismatch and compounding errors. DAgger is an iterative algorithm that addresses these issues by aggregating training data from both the expert and novice policies, but does not consider the impact of sa…

Cited by 127SourceScholar
2019

HG-DAgger: Interactive Imitation Learning with Human Experts

ICRA 2019poster

Imitation learning has proven to be useful for many real-world problems, but approaches such as behavioral cloning suffer from data mismatch and compounding error issues. One attempt to address these limitations is the DAgger algorithm, which uses the state distribution induced by the novice to samp…

Cited by 251SourceScholar
2019

Limiting Extrapolation in Linear Approximate Value Iteration

NeurIPS 2019poster

We study linear approximate value iteration (LAVI) with a generative model. While linear models may accurately represent the optimal value function using a few parameters, several empirical and theoretical studies show the combination of least-squares projection with the Bellman operator may be expa…

Cited by 38SourcePDFScholar
2019

Simulating Emergent Properties of Human Driving Behavior Using Multi-Agent Reward Augmented Imitation Learning

ICRA 2019poster

Recent developments in multi-agent imitation learning have shown promising results for modeling the behavior of human drivers. However, it is challenging to capture emergent traffic behaviors that are observed in real-world datasets. Such behaviors arise due to the many local interactions between ag…

Cited by 72SourcecodeScholar
2018

Deep Dynamical Modeling and Control of Unsteady Fluid Flows

NeurIPS 2018poster

The design of flow control systems remains a challenge due to the nonlinear nature of the equations that govern fluid flow. However, recent advances in computational fluid dynamics (CFD) have enabled the simulation of complex fluid flows with high accuracy, opening the possibility of using learning-…

2018

Gaussian Process Dynamic Programming for Optimizing Ungrounded Haptic Guidance

IROS 2018poster

Adapting robot actions to human motions can make human-robot interactions (HRI) more effective. Here, we aim to optimize guidance from haptic devices based on a user's response to produce better task performance. We used Gaussian processes to model the motions a human user made in response to applie…

Cited by 5SourceScholar
2018

Improving Offline Value-Function Approximations for POMDPs by Reducing Discount Factors

IROS 2018poster

A common solution criterion for partially observable Markov decision processes (POMDPs) is to maximize the expected sum of exponentially discounted rewards, for which a variety of approximate methods have been proposed. Those that plan in the belief space typically provide tighter performance guaran…

Cited by 8SourceScholar
2018

Multi-Agent Imitation Learning for Driving Simulation

IROS 2018poster

Simulation is an appealing option for validating the safety of autonomous vehicles. Generative Adversarial Imitation Learning (GAIL) has recently been shown to learn representative human driver models. These human driver models were learned through training in single-agent environments, but they hav…

Cited by 143SourcecodeScholar
2018

People as Sensors: Imputing Maps from Human Actions

IROS 2018poster

Despite growing attention in autonomy, there are still many open problems, including how autonomous vehicles will interact and communicate with other agents, such as human drivers and pedestrians. Unlike most approaches that focus on pedestrian detection and planning for collision avoidance, this pa…

Cited by 36SourceScholar
2018

Pseudo-bearing Measurements for Improved Localization of Radio Sources with Multirotor UAVs

ICRA 2018poster

Localizing radio frequency (RF) sources is an important application for unmanned aerial vehicles (UAVs), Localization is often carried out by estimating bearing to an RF source, which can be achieved by rotating a directional antenna in place. Multirotor UAVs are well-suited for this sensing modalit…

Cited by 21SourceScholar
2018

Scalable Decision Making with Sensor Occlusions for Autonomous Driving

ICRA 2018poster

Autonomous driving in urban areas requires avoiding other road users with only partial observability of the environment. Observations are only partial because obstacles can occlude the field of view of the sensors. The problem of robust and efficient navigation under uncertainty can be framed as a p…

Cited by 86SourceScholar
2016

Optimized and trusted collision avoidance for unmanned aerial vehicles using approximate dynamic programming

ICRA 2016

Safely integrating unmanned aerial vehicles into civil airspace is contingent upon development of a trustworthy collision avoidance system. This paper proposes an approach whereby a parameterized resolution logic that is considered trusted for a given range of its parameters is adaptively tuned onli

Cited by 21SourceScholar