← Search

Lihan Zha

7 accepted papers

2026

Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting

ICLR 2026poster

Fine-tuning vision-language models (VLMs) on robot teleoperation data to create vision-language-action (VLA) models is a promising paradigm for training generalist policies, but it suffers from a fundamental tradeoff: learning to produce actions often diminishes the VLM's foundational reasoning and…

Cited by 0SourceScholar
2026

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning

RSS 2026poster

Robotic systems that interact with the physical world must reason about kinematic and dynamic constraints imposed by their own embodiment, their environment, and the task at hand. We introduce KinDER, a benchmark for Kinematic and Dynamic Embodied Reasoning that targets physical reasoning challenges…

Cited by 0SourceScholar
2026

LAP: Language-Action Pre-training Enables Zero-Shot Cross-Embodiment Transfer

RSS 2026poster

A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision–Language–Action models (VLAs) remain tightly coupled to their training embodiments and…

Cited by 0SourceScholar
2026

Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators

ICRA 2026poster

Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often eval…

2025

WoMAP: World Models For Embodied Open-Vocabulary Object Localization

CoRL 2025poster

Active object localization remains a critical challenge for robots, requiring efficient exploration of partially observable environments. However, state-of-the-art robot policies either struggle to generalize beyond demonstration datasets (e.g., imitation learning methods) or fail to generate physic…

Cited by 0SourceScholar
2024

Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections

ICRA 2024poster

Today’s robot policies exhibit subpar performance when faced with the challenge of generalizing to novel environments. Human corrective feedback is a crucial form of guidance to enable such generalization. However, adapting to and learning from online human corrections is a non-trivial endeavor: not…

Cited by 43SourcecodeScholar
2024

DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment

IROS 2024poster

Large language models (LLMs) encode a vast amount of semantic knowledge and possess remarkable understanding and reasoning capabilities. Previous work has explored how to ground LLMs in robotic tasks to generate feasible and executable textual plans. However, low-level execution in the physical worl…

Cited by 37SourceScholar