← Search

Haojie Huang

12 accepted papers

2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

ICML 2026poster

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improvement. However, existing embodied benchmarks fail to provide actionable insights because they focus on task-level evaluation rather than discovering capability bottlenecks. To address t…

Cited by 0SourceScholar
2026

Generalizable Hierarchical Skill Learning via Object-Centric Representation

RA-L 2026

We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that improves policy generalization and sample efficiency in robot manipulation. One core idea of GSL is to use object-centric skills as an interface that bridges the high-level vision-lang

Cited by 3SourceScholar
2025

3D Equivariant Visuomotor Policy Learning via Spherical Projection

NeurIPS 2025spotlight

Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused primarily on point cloud inputs generated by multiple cameras fixed in the workspace. This type of point cloud input is not…

Cited by 0SourceScholar
2025

Learning Efficient and Robust Language-Conditioned Manipulation Using Textual-Visual Relevancy and Equivariant Language Mapping

RA-L 2025

Controlling robots through natural language is pivotal for enhancing human-robot collaboration and synthesizing complex robot behaviors. Recent works that are trained on large robot datasets show impressive generalization abilities. However, such pretrained methods are (1) often fragile to unseen sc

Cited by 7SourcecodeScholar
2025

Match Policy: A Simple Pipeline from Point Cloud Registration to Manipulation Policies

ICRA 2025

Many manipulation tasks require the robot to rearrange objects relative to one another. Such tasks can be described as a sequence of relative poses between parts of a set of rigid bodies. In this work, we propose Match Policy, a simple but novel pipeline for solving high-precision pick and place tas

Cited by 5SourcecodeScholar
2024

Equivariant Diffusion Policy

CoRL 2024poster

Recent work has shown diffusion models are an effective approach to learning the multimodal distributions arising from demonstration data in behavior cloning. However, a drawback of this approach is the need to learn a denoising function, which is significantly more complex than learning an explicit…

Cited by 26SourcecodeScholar
2024

Fourier Transporter: Bi-Equivariant Robotic Manipulation in 3D

ICLR 2024poster

Many complex robotic manipulation tasks can be decomposed as a sequence of pick and place actions. Training a robotic agent to learn this sequence over many different starting conditions typically requires many iterations or demonstrations, especially in 3D environments. In this work, we propose Fou…

Cited by 24SourcePDFScholar
2024

IMAGINATION POLICY: Using Generative Point Cloud Models for Learning Manipulation Policies

CoRL 2024poster

Humans can imagine goal states during planning and perform actions to match those goals. In this work, we propose IMAGINATION POLICY, a novel multi-task key-frame policy network for solving high-precision pick and place tasks. Instead of learning actions directly, IMAGINATION POLICY generates point…

Cited by 7SourceScholar
2024

OrbitGrasp: SE(3)-Equivariant Grasp Learning

CoRL 2024poster

While grasp detection is an important part of any robotic manipulation pipeline, reliable and accurate grasp detection in $\\mathrm{SE}(3)$ remains a research challenge. Many robotics applications in unstructured environments such as the home or warehouse would benefit a lot from better grasp perfor…

Cited by 13SourcecodeScholar
2024

ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter

CoRL 2024poster

Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-language grasping system that makes use of GPT-4o's advanced contextual reasoning for grasping strategies. ThinkGrasp can…

Cited by 14SourcecodeScholar
2023

Edge Grasp Network: A Graph-Based SE(3)-invariant Approach to Grasp Detection

ICRA 2023poster

Given point cloud input, the problem of 6-DoF grasp pose detection is to identify a set of hand poses in SE(3) from which an object can be successfully grasped. This important problem has many practical applications. Here we propose a novel method and neural network model that enables better grasp s…

Cited by 39SourcecodeScholar