← Search

Vaibhav Saxena

6 accepted papers

2026

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning

RSS 2026poster

Robotic systems that interact with the physical world must reason about kinematic and dynamic constraints imposed by their own embodiment, their environment, and the task at hand. We introduce KinDER, a benchmark for Kinematic and Dynamic Embodied Reasoning that targets physical reasoning challenges…

Cited by 0SourceScholar
2025

Towards Robust Knowledge Representations in Multilingual LLMs for Equivalence and Inheritance based Consistent Reasoning

NAACL 2025long

Reasoning and linguistic skills form the cornerstone of human intelligence, facilitating problem-solving and decision-making. Recent advances in Large Language Models (LLMs) have led to impressive linguistic capabilities and emergent reasoning behaviors, fueling widespread adoption across applicatio…

Cited by 2SourcePDFScholar
2025

What Matters in Learning from Large-Scale Datasets for Robot Manipulation

ICLR 2025poster

Imitation learning from large multi-task demonstration datasets has emerged as a promising path for building generally-capable robots. As a result, 1000s of hours have been spent on building such large-scale datasets around the globe. Despite the continuous growth of such efforts, we still lack a sy…

Cited by 3SourcePDFScholar
2024

MimicTouch: Leveraging Multi-modal Human Tactile Demonstrations for Contact-rich Manipulation

CoRL 2024poster

Tactile sensing is critical to fine-grained, contact-rich manipulation tasks, such as insertion and assembly. Prior research has shown the possibility of learning tactile-guided policy from teleoperated demonstration data. However, to provide the demonstration, human users often rely on visual feedb…

Cited by 16SourceScholar
2023

Generalizable Pose Estimation Using Implicit Scene Representations

ICRA 2023poster

6-DoF pose estimation is an essential component of robotic manipulation pipelines. However, it usually suffers from a lack of generalization to new instances and object types. Most widely used methods learn to infer the object pose in a discriminative setup where the model filters useful information…

Cited by 13SourcecodeScholar