← Search

Masha Itkina

23 accepted papers

2026

Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison

RSS 2026poster

Generalist robot manipulation policies are becoming increasingly capable, but are limited in evaluation to a small number of hardware rollouts. This strong resource constraint in real-world testing necessitates both more informative performance measures and reliable and efficient evaluation procedur…

Cited by 0SourceScholar
2026

Impact of Different Failures on a Robot’s Perceived Reliability

ICRA 2026poster

Robots fail, potentially leading to a loss in the robot’s perceived reliability (PR), a measure correlated with trustworthiness. In this study we examine how various kinds of failures affect the PR of the robot differently, and how this measure recovers without explicit social repair actions by the …

2026

Using Non-Expert Data to Robustify Imitation Learning Via Offline Reinforcement Learning

ICRA 2026poster

Imitation learning has proven effective for training robots to perform complex tasks from expert human demonstrations. However, it remains limited by its reliance on high-quality, task-specific data, restricting adaptability to the diverse range of real-world object configurations and scenarios. In …

2025

CUPID: Curating Data your Robot Loves with Influence Functions

CoRL 2025poster

In robot imitation learning, policy performance is tightly coupled with the quality and composition of the demonstration data. Yet, developing a precise understanding of how individual demonstrations contribute to downstream outcomes—such as closed-loop task success or failure—remains a persistent c…

Cited by 0SourceScholar
2025

Can We Detect Failures Without Failure Data? Uncertainty-Aware Runtime Failure Detection for Imitation Learning Policies

RSS 2025poster

Recent years have witnessed impressive robotic manipulation systems driven by advances in imitation learning and generative modeling, such as diffusion- and flow-based approaches. As robot policy performance increases, so does the complexity and time horizon of achievable tasks, inducing unexpected…

Cited by 1PDFScholar
2025

GHIL-Glue: Hierarchical Control with Filtered Subgoal Images

ICRA 2025

Image and video generative models that are pretrained on Internet-scale data can greatly increase the generalization capacity of robot learning systems. These models can function as high-level planners, generating intermediate sub-goals for low-level goal-conditioned policies to reach. However, the

Cited by 9SourcecodeScholar
2025

Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping

RSS 2025poster

Imitation learning has enabled robots to perform complex, long-horizon tasks in challenging dexterous manipulation settings. As new methods are developed, they must be rigorously evaluated and compared against corresponding baselines through repeated evaluation trials. However, policy comparison is…

Cited by 1PDFScholar
2025

SAFE: Multitask Failure Detection for Vision-Language-Action Models

NeurIPS 2025poster

While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on novel tasks out of the box. To allow these policies to safely interact with their environments, we need a failure detector…

Cited by 0SourcecodeScholar
2025

STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation

NeurIPS 2025spotlight

Off-policy evaluation (OPE) estimates the performance of a target policy using offline data collected from a behavior policy, and is crucial in domains such as robotics or healthcare where direct interaction with the environment is costly or unsafe. Existing OPE methods are ineffective for high-dime…

Cited by 0SourceScholar
2025

Self-supervised Multi-future Occupancy Forecasting for Autonomous Driving

RSS 2025poster

Environment prediction frameworks are critical for the safe navigation of autonomous vehicles (AVs) in dynamic settings. LiDAR-generated occupancy grid maps (L-OGMs) offer a robust bird’s-eye view scene representation, enabling self-supervised joint scene predictions while exhibiting resilience to p…

Cited by 3PDFScholar
2024

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

RSS 2024poster

The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistica…

Cited by 216SourcePDFScholar
2024

Explore until Confident: Efficient Exploration for Embodied Question Answering

RSS 2024poster

We consider the problem of Embodied Question Answering (EQA), which refers to settings where an embodied agent such as a robot needs to actively explore an environment to gather information until it is confident about the answer to a question. In this work, we leverage the strong semantic reasoning…

Cited by 39SourcePDFScholar
2024

How Generalizable is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation

RA-L 2024

With the rise of stochastic generative models in robot policy learning, end-to-end visuomotor policies are increasingly successful at solving complex tasks by learning from human demonstrations. Nevertheless, since real-world evaluation costs afford users only a small number of policy rollouts, it r

Cited by 15SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2023

Occlusion-Aware Crowd Navigation Using People as Sensors

ICRA 2023poster

Autonomous navigation in crowded spaces poses a challenge for mobile robots due to the highly dynamic, partially observable environment. Occlusions are highly prevalent in such settings due to a limited sensor field of view and obstructing human agents. Previous work has shown that observed interact…

Cited by 21SourcecodeScholar
2022

How Do We Fail? Stress Testing Perception in Autonomous Vehicles

IROS 2022poster

Autonomous vehicles (AVs) rely on environment perception and behavior prediction to reason about agents in their surroundings. These perception systems must be robust to adverse weather such as rain, fog, and snow. However, validation of these systems is challenging due to their complexity and depen…

Cited by 21SourcecodeScholar
2022

Multi-Agent Variational Occlusion Inference Using People as Sensors

ICRA 2022poster

Autonomous vehicles must reason about spatial occlusions in urban environments to ensure safety without being overly cautious. Prior work explored occlusion inference from observed social behaviors of road agents, hence treating people as sensors. Inferring occupancy from agent behaviors is an inher…

Cited by 27SourcecodeScholar
2021

Double-Prong ConvLSTM for Spatiotemporal Occupancy Prediction in Dynamic Environments

ICRA 2021poster

Predicting the future occupancy state of an environment is important to enable informed decisions for autonomous vehicles. Common challenges in occupancy prediction include vanishing dynamic objects and blurred predictions, especially for long prediction horizons. In this work, we propose a double-p…

Cited by 26SourcecodeScholar
2021

Evidential Softmax for Sparse Multimodal Distributions in Deep Generative Models

NeurIPS 2021poster

Many applications of generative models rely on the marginalization of their high-dimensional output probability distributions. Normalization functions that yield sparse probability distributions can make exact marginalization more computationally tractable. However, sparse normalization functions us…

2020

Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders

NeurIPS 2020poster

Discrete latent spaces in variational autoencoders have been shown to effectively capture the data distribution for many real-world problems such as natural language understanding, human intent prediction, and visual scene representation. However, discrete latent spaces need to be sufficiently large…