← Search

Muzhi Han

9 accepted papers

2026

M3Bench: Benchmarking Whole-Body Motion Generation for Mobile Manipulation in 3D Scenes

ICRA 2026poster

We propose M3Bench, a new benchmark for whole-body motion generation in mobile manipulation tasks. Given a 3D scene context, M3Bench requires an embodied agent to reason about its configuration, environmental constraints, and task objectives to generate coordinated whole-body motion trajectories for…

2025

Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V

ICRA 2025

Autonomous robot navigation and manipulation in open environments require reasoning and replanning with closed-loop feedback. In this work, we present COME-robot, the first closed-loop robotic system utilizing the GPT-4V vision-language foundation model for open-ended reasoning and adaptive planning

Cited by 62SourceScholar
2025

M${}{3}$Bench: Benchmarking Whole-Body Motion Generation for Mobile Manipulation in 3D Scenes

RA-L 2025

We propose M <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">${}^{3}$</tex-math></inline-formula> Bench, a new benchmark for whole-body motion generation in mobile manipulation tasks. Given a 3D scene context, M <in

Cited by 4SourceScholar
2024

Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations

IROS 2024poster

Autonomous robotic systems capable of learning novel manipulation tasks are poised to transform industries from manufacturing to service automation. However, current methods (e.g., VIP and R3M) still face significant hurdles, notably the domain gap among robotic embodiments and the sparsity of succe…

Cited by 15SourcecodeScholar
2024

INTERPRET: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning

RSS 2024poster

Learning abstract state representations and knowledge is crucial for long-horizon robot planning. We present InterPreT, an LLM-powered framework for robots to learn symbolic predicates from language feedback of human non-experts during embodied interaction. The learned predicates provide relational…

2024

LLM3: Large Language Model-based Task and Motion Planning with Motion Failure Reasoning

IROS 2024poster

Conventional Task and Motion Planning (TAMP) approaches rely on manually designed interfaces connecting symbolic task planning with continuous motion generation. These domain-specific and labor-intensive modules are limited in addressing emerging tasks in real-world settings. Here, we present LLM3,…

Cited by 45SourcecodeScholar
2023

Learning a Causal Transition Model for Object Cutting

IROS 2023poster

Cutting objects into desired fragments is challenging for robots due to the spatially unstructured nature of fragments and the complex one-to-many object fragmentation caused by actions. We present a novel approach to model object fragmentation using an attributed stochastic grammar. This grammar ab…

Cited by 2SourceScholar
2023

Part-level Scene Reconstruction Affords Robot Interaction

IROS 2023poster

Existing methods for reconstructing interactive scenes primarily focus on replacing reconstructed objects with CAD models retrieved from a limited database, resulting in significant discrepancies between the reconstructed and observed scenes. To address this issue, our work introduces a part-level r…

Cited by 9SourceScholar
2021

Reconstructing Interactive 3D Scenes by Panoptic Mapping and CAD Model Alignments

ICRA 2021poster

In this paper, we rethink the problem of scene reconstruction from an embodied agent’s perspective: While the classic view focuses on the reconstruction accuracy, our new perspective emphasizes the underlying functions and constraints such that the reconstructed scenes provide actionable information…

Cited by 32SourcecodeScholar