← Search

Apurva Badithela

5 accepted papers

2026

Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison

RSS 2026poster

Generalist robot manipulation policies are becoming increasingly capable, but are limited in evaluation to a small number of hardware rollouts. This strong resource constraint in real-world testing necessitates both more informative performance measures and reliable and efficient evaluation procedur…

Cited by 0SourceScholar
2026

Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators

ICRA 2026poster

Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often eval…

2025

Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping

RSS 2025poster

Imitation learning has enabled robots to perform complex, long-horizon tasks in challenging dexterous manipulation settings. As new methods are developed, they must be rigorously evaluated and compared against corresponding baselines through repeated evaluation trials. However, policy comparison is…

Cited by 1PDFScholar
2023

Evaluation Metrics of Object Detection for Quantitative System-Level Analysis of Safety-Critical Autonomous Systems

IROS 2023poster

This paper proposes two metrics for evaluating learned object detection models: the proposition-labeled and distance-parametrized confusion matrices. These metrics are leveraged to quantitatively analyze the system with respect to its system-level formal specifications via probabilistic model checki…

Cited by 8SourceScholar
2023

Synthesizing Reactive Test Environments for Autonomous Systems: Testing Reach-Avoid Specifications with Multi-Commodity Flows

ICRA 2023poster

We study automated test generation for testing discrete decision-making modules in autonomous systems. Linear temporal logic is used to encode the system specification - requirements of the system under test - and the test specification, which is unknown to the system and describes the desired test…

Cited by 7SourceScholar