← Search

Anirudha Majumdar

39 accepted papers

2026

Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting

ICLR 2026poster

Fine-tuning vision-language models (VLMs) on robot teleoperation data to create vision-language-action (VLA) models is a promising paradigm for training generalist policies, but it suffers from a fundamental tradeoff: learning to produce actions often diminishes the VLM's foundational reasoning and…

Cited by 0SourceScholar
2026

Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison

RSS 2026poster

Generalist robot manipulation policies are becoming increasingly capable, but are limited in evaluation to a small number of hardware rollouts. This strong resource constraint in real-world testing necessitates both more informative performance measures and reliable and efficient evaluation procedur…

Cited by 0SourceScholar
2026

LAP: Language-Action Pre-training Enables Zero-Shot Cross-Embodiment Transfer

RSS 2026poster

A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision–Language–Action models (VLAs) remain tightly coupled to their training embodiments and…

Cited by 0SourceScholar
2026

Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators

ICRA 2026poster

Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often eval…

2025

Diffusion Policy Policy Optimization

ICLR 2025poster

We introduce Diffusion Policy Policy Optimization, DPPO, an algorithmic framework including best practices for fine-tuning diffusion-based policies (e.g. Diffusion Policy) in continuous control and robot learning tasks using the policy gradient (PG) method from reinforcement learning (RL). PG method…

Cited by 270SourcePDFScholar
2025

Generating Robot Constitutions & Benchmarks for Semantic Safety

CoRL 2025poster

Large vision and language models are being increasingly deployed on real robots, leading to an immediate need for ensuring robot safety under AI-control. In this paper, we develop the ASIMOV Benchmark — a collection of large-scale semantic safety datasets grounded in real-world visual scenes and hum…

Cited by 0SourceScholar
2025

Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping

RSS 2025poster

Imitation learning has enabled robots to perform complex, long-horizon tasks in challenging dexterous manipulation settings. As new methods are developed, they must be rigorously evaluated and compared against corresponding baselines through repeated evaluation trials. However, policy comparison is…

Cited by 1PDFScholar
2025

Predictive Red Teaming: Breaking Policies Without Breaking Robots

CoRL 2025poster

Visuomotor policies trained via imitation learning are capable of performing challenging manipulation tasks, but are often extremely brittle to lighting, visual distractors, and object locations. These vulnerabilities can depend unpredictably on the specifics of training, and are challenging to expo…

Cited by 0SourceScholar
2025

Run-time Observation Interventions Make Vision-Language-Action Models More Visually Robust

ICRA 2025

Vision-language-action (VLA) models trained on large-scale internet data and robot demonstrations have the potential to serve as generalist robot policies. However, despite their large-scale training, VLAs are often brittle to task-irrelevant visual details such as distractor objects or background c

Cited by 24SourcecodeScholar
2025

SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps

CoRL 2025poster

We present SIREN for registration of multi-robot Gaussian Splatting (GSplat) maps, with zero access to camera poses, images, and inter-map transforms for initialization or fusion of local submaps. To realize these capabilities, SIREN harnesses the versatility and robustness of semantics in three cri…

Cited by 0SourceScholar
2025

WoMAP: World Models For Embodied Open-Vocabulary Object Localization

CoRL 2025poster

Active object localization remains a critical challenge for robots, requiring efficient exploration of partially observable environments. However, state-of-the-art robot policies either struggle to generalize beyond demonstration datasets (e.g., imitation learning methods) or fail to generate physic…

Cited by 0SourceScholar
2024

Explore until Confident: Efficient Exploration for Embodied Question Answering

RSS 2024poster

We consider the problem of Embodied Question Answering (EQA), which refers to settings where an embodied agent such as a robot needs to actively explore an environment to gather information until it is confident about the answer to a question. In this work, we leverage the strong semantic reasoning…

Cited by 39SourcePDFScholar
2024

Perceive With Confidence: Statistical Safety Assurances for Navigation with Learning-Based Perception

CoRL 2024poster

Rapid advances in perception have enabled large pre-trained models to be used out of the box for transforming high-dimensional, noisy, and partial observations of the world into rich occupancy representations. However, the reliability of these models and consequently their safe integration onto robo…

Cited by 8SourceScholar
2024

Physically Grounded Vision-Language Models for Robotic Manipulation

ICRA 2024poster

Recent advances in vision-language models (VLMs) have led to improved performance on tasks such as visual question answering and image captioning. Consequently, these models are now well-positioned to reason about the physical world, particularly within domains such as robotic manipulation. However,…

Cited by 131SourceScholar
2024

Privacy-Preserving Map-Free Exploration for Confirming the Absence of a Radioactive Source

IROS 2024poster

Performing an inspection task while maintaining the privacy of the inspected site is a challenging balancing act. In this work, we are motivated by the future of nuclear arms control verification, which requires both a high level of privacy and guaranteed correctness. For scenarios with limitations…

Cited by 0SourcecodeScholar
2024

Risk-Calibrated Human-Robot Interaction via Set-Valued Intent Prediction

RSS 2024poster

Tasks where robots must anticipate human intent, such as navigating around a cluttered home or sorting everyday items, are challenging because they exhibit a wide range of valid actions that lead to similar outcomes. Moreover, zero-shot cooperation between human-robot partners is an especially chall…

Cited by 6SourcePDFScholar
2024

Sim-to-Lab-to-Real: Safe Reinforcement Learning with Shielding and Generalization Guarantees (Abstract Reprint)

AAAI 2024technical

Safety is a critical component of autonomous systems and remains a challenge for learning-based policies to be utilized in the real world. In particular, policies learned using reinforcement learning often fail to generalize to novel environments due to unsafe behavior. In this paper, we propose Sim…

Cited by 1SourcePDFScholar
2023

AdaptSim: Task-Driven Simulation Adaptation for Sim-to-Real Transfer

CoRL 2023poster

Simulation parameter settings such as contact models and object geometry approximations are critical to training robust manipulation policies capable of transferring from simulation to real-world deployment. There is often an irreducible gap between simulation and reality: attempting to match the dy…

Cited by 16SourceScholar
2023

FlowDrone: Wind Estimation and Gust Rejection on UAVs Using Fast-Response Hot-Wire Flow Sensors

ICRA 2023poster

Unmanned aerial vehicles (UAVs) are finding use in applications that place increasing emphasis on robustness to external disturbances including extreme wind. However, traditional multirotor UAV platforms do not directly sense wind; conventional flow sensors are too slow, insensitive, or bulky for wi…

Cited by 19SourceScholar
2023

Online Learning for Obstacle Avoidance

CoRL 2023poster

We approach the fundamental problem of obstacle avoidance for robotic systems via the lens of online learning. In contrast to prior work that either assumes worst-case realizations of uncertainty in the environment or a stationary stochastic model of uncertainty, we propose a method that is efficien…

Cited by 3SourceScholar
2023

PAC-Bayes Generalization Certificates for Learned Inductive Conformal Prediction

NeurIPS 2023poster

Inductive Conformal Prediction (ICP) provides a practical and effective approach for equipping deep learning models with uncertainty estimates in the form of set-valued predictions which are guaranteed to contain the ground truth with high probability. Despite the appeal of this coverage guarantee,…

Cited by 9SourcePDFScholar
2023

Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners

CoRL 2023oral

Large language models (LLMs) exhibit a wide range of promising capabilities --- from step-by-step planning to commonsense reasoning --- that may provide utility for robots, but remain prone to confidently hallucinated predictions. In this work, we present KnowNo, a framework for measuring and aligni…

Cited by 248SourceScholar
2023

Switching Attention in Time-Varying Environments via Bayesian Inference of Abstractions

ICRA 2023poster

Motivated by the goal of endowing robots with a means for focusing attention in order to operate reliably in complex, uncertain, and time-varying environments, we consider how a robot can (i) determine which portions of its environment to pay attention to at any given point in time, (ii) infer chang…

Cited by 1SourceScholar
2022

Fundamental Performance Limits for Sensor-Based Robot Control and Policy Learning

RSS 2022poster

Our goal is to develop theory and algorithms for establishing fundamental limits on performance for a given task imposed by a robot's sensors. In order to achieve this, we define a quantity that captures the amount of task-relevant information provided by a sensor. Using a novel version of the gener…

2022

Leveraging Language for Accelerated Learning of Tool Manipulation

CoRL 2022poster

Robust and generalized tool manipulation requires an understanding of the properties and affordances of different tools. We investigate whether linguistic information about a tool (e.g., its geometry, common uses) can help control policies adapt faster to new tools for a given task. We obtain divers…

Cited by 53SourceScholar
2022

Robust Control Under Uncertainty via Bounded Rationality and Differential Privacy

ICRA 2022poster

The rapid development of affordable and compact high-fidelity sensors (e.g., cameras and LIDAR) allows robots to construct detailed estimates of their states and environments. However, the availability of such rich sensor information introduces two challenges: (i) the lack of analytic sensing models…

Cited by 11SourcecodeScholar
2022

Stronger Generalization Guarantees for Robot Learning by Combining Generative Models and Real-World Data

ICRA 2022poster

We are motivated by the problem of learning policies for robotic systems with rich sensory inputs (e.g., vision) in a manner that allows us to guarantee generalization to environments unseen during training. We provide a framework for providing such generalization guarantees by leveraging a finite d…

Cited by 2SourceScholar
2021

A Regret Minimization Approach to Iterative Learning Control

ICML 2021spotlight

We consider the setting of iterative learning control, or model-based policy learning in the presence of uncertain, time-varying dynamics. In this setting, we propose a new performance metric, planning regret, which replaces the standard stochastic uncertainty assumptions with worst case regret. Bas…

2021

Generalization Bounds for Meta-Learning via PAC-Bayes and Uniform Stability

NeurIPS 2021poster

We are motivated by the problem of providing strong generalization guarantees in the context of meta-learning. Existing generalization bounds are either challenging to evaluate or provide vacuous guarantees in even relatively simple settings. We derive a probably approximately correct (PAC) bound fo…

2021

Task-Driven Out-of-Distribution Detection with Statistical Guarantees for Robot Learning

CoRL 2021poster

Our goal is to perform out-of-distribution (OOD) detection, i.e., to detect when a robot is operating in environments that are drawn from a different distribution than the environments used to train the robot. We leverage Probably Approximately Correct (PAC)-Bayes theory in order to train a policy w…

Cited by 32SourceScholar
2018

PAC-Bayes Control: Synthesizing Controllers that Provably Generalize to Novel Environments

CoRL 2018

Our goal is to synthesize controllers for robots that provably generalize well to novel environments given a dataset of example environments. The key technical idea behind our approach is to leverage tools from generalization theory in machine learning by exploiting a precise analogy (which we prese

2017

Risk-sensitive Inverse Reinforcement Learning via Coherent Risk Models

RSS 2017poster

The literature on Inverse Reinforcement Learning (IRL) typically assumes that humans take actions in order to minimize the expected value of a cost function, i.e., that humans are risk neutral. Yet, in practice, humans are often far from being risk neutral. To fill this gap, the objective of this pa…

Cited by 87SourcePDFScholar
2017

Robust online motion planning via contraction theory and convex optimization

ICRA 2017poster

We present a framework for online generation of robust motion plans for robotic systems with nonlinear dynamics subject to bounded disturbances, control constraints, and online state constraints such as obstacles. In an offline phase, one computes the structure of a feedback controller that can be e…

Cited by 237SourceScholar