← Search

Huang Huang

34 accepted papers

2026

CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

ICML 2026poster

“Code-as-Policy” considers how executable code can complement data-intensive Vision-LanguageAction (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaPX, an open-access framework for systematically studying Code-as-Policy ag…

Cited by 0SourcecodeScholar
2026

Cross-Embodiment Robot Foundation World Models with Latent Actions

ICML 2026poster

The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce a Latent Action Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse e…

Cited by 0SourceScholar
2026

MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation

ICLR 2026poster

Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This challenge intensifies for multi-step bimanual mobile manipulation, where humans must teleoperate both the mobile base and…

Cited by 0SourcecodeScholar
2025

Blox-Net: Generative Design-for-Robot-Assembly Using VLM Supervision, Physics Simulation, and a Robot with Reset

ICRA 2025

Generative AI systems have shown impressive capabilities in creating text, code, and images. Inspired by the importance of research in industrial Design for Assembly, we introduce a novel problem: Generative Design-for-RobotAssembly (GDfRA). The task is to generate an assembly based on a natural lan

Cited by 16SourceScholar
2025

ICRT: In-Context Imitation Learning via Next-Token Prediction

ICRA 2025

In-context imitation learning is the capability to perform novel tasks when prompted with task demonstration examples. In-Context Robot Transformer (ICRT) is a causal transformer that performs autoregressive prediction on sensorimotor trajectories, which include images, proprioceptive states, and ac

Cited by 53SourceScholar
2025

OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction

ICML 2025poster

Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained vision-language models (VLMs) as visual and language features are independently fed into downstream policies, degrading the p…

2025

Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding

ICCV 2025poster

Whether snipping with scissors or opening a box, humans can quickly understand the 3D configurations of familiar objects. For novel objects, we can resort to long-form inspection to build intuition. The more we observe the object, the better we get at predicting its 3D state immediately. Existing sy…

Cited by 0SourcePDFScholar
2025

Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware

CoRL 2025oral

Scaling robot learning requires vast and diverse datasets. Yet the prevailing data collection paradigm—human teleoperation—remains costly and constrained by manual effort and physical robot access. We introduce Real2Render2Real (R2R2R), a novel approach for generating robot training data without rel…

Cited by 0SourceScholar
2025

Robo-DM: Data Management for Large Robot Datasets

ICRA 2025

Recent results suggest that very large datasets of teleoperated robot demonstrations can be used to train transformer-based models that have the potential to generalize to new scenes, robots, and tasks. However, curating, distributing, and loading large datasets of robot trajectories, which typicall

Cited by 1SourcecodeScholar
2025

Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion

ACL 2025long

This paper addresses the critical need for democratizing large language models (LLM) in the Arab world, a region that has seen slower progress in developing models comparable to state-of-the-art offerings like GPT-4 or GPT-3.5, due to a predominant focus on mainstream languages (e.g., English and Ch…

2024

A Touch, Vision, and Language Dataset for Multimodal Alignment

ICML 2024oral

Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactile data and the complexity of aligning tactile readings with both visual observat…

2024

AceGPT, Localizing Large Language Models in Arabic

NAACL 2024long

This paper is devoted to the development of a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models. Significant concerns emerge when addressing cultural sensitivity and local values. T…

2024

Alignment at Pre-training! Towards Native Alignment for Arabic LLMs

NeurIPS 2024poster

The alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `\textit{post alignment}'. We argue that alignment…

2024

Conformal Policy Learning for Sensorimotor Control under Distribution Shifts

ICRA 2024poster

This paper focuses on the problem of detecting and reacting to changes in the distribution of a sensorimotor controller’s observables. The key idea is the design of policies that can take conformal quantiles as input, to detect distribution shifts with formal statistical guarantees, which we define…

Cited by 5SourceScholar
2024

DiffusionSeeder: Seeding Motion Optimization with Diffusion for Rapid Motion Planning

CoRL 2024poster

Running optimization across many parallel seeds leveraging GPU compute [2] have relaxed the need for a good initialization, but this can fail if the problem is highly non-convex as all seeds could get stuck in local minima. One such setting is collision-free motion optimization for robot manipulatio…

Cited by 33SourceScholar
2024

Lifelong LERF: Local 3D Semantic Inventory Monitoring Using FogROS2

ICRA 2024poster

Inventory monitoring in homes, factories, and retail stores relies on maintaining data despite objects being swapped, added, removed, or moved. We introduce Lifelong LERF, a method that allows a mobile robot with minimal compute to jointly optimize a dense language and geometric representation of it…

Cited by 6SourceScholar
2024

Manipulator as a Tail: Promoting Dynamic Stability for Legged Locomotion

ICRA 2024poster

For locomotion, is an arm on a legged robot a liability or an asset for locomotion? Biological systems evolved additional limbs beyond legs that facilitates postural control. This work shows how a manipulator can be an asset for legged locomotion at high speeds or under external perturbations, where…

Cited by 5SourceScholar
2024

Orbit-Surgical: An Open-Simulation Framework for Learning Surgical Augmented Dexterity

ICRA 2024poster

Physics-based simulations have accelerated progress in robot learning for driving, manipulation, and locomotion. Yet, a fast, accurate, and robust surgical simulation environment remains a challenge. In this paper, we present Orbit-Surgical, a physics-based surgical robot simulation framework with p…

Cited by 17SourcecodeScholar
2023

Automating Vascular Shunt Insertion with the dVRK Surgical Robot

ICRA 2023poster

Vascular shunt insertion is a fundamental surgical procedure used to temporarily restore blood flow to tissues. It is often performed in the field after major trauma. We formulate a problem of automated vascular shunt insertion and propose a pipeline to perform Automated Vascular Shunt Insertion (AV…

Cited by 11SourceScholar
2023

Grasp Stability Assessment Through Attention-Guided Cross-Modality Fusion and Transfer Learning

IROS 2023poster

Extensive research has been conducted on assessing grasp stability, a crucial prerequisite for achieving optimal grasping strategies, including the minimum force grasping policy. However, existing works employ basic feature-level fusion techniques to combine visual and tactile modalities, resulting…

Cited by 9SourceScholar
2023

Safe Self-Supervised Learning in Real of Visuo-Tactile Feedback Policies for Industrial Insertion

ICRA 2023poster

Industrial insertion tasks are often performed repetitively with parts that are subject to tight tolerances and prone to breakage. Learning an industrial insertion policy in real is challenging as the collision between the parts and the environment can cause slippage or breakage of the part. In this…

Cited by 22SourceScholar
2023

Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features

RSS 2023poster

Humans make extensive use of vision and touch as complementary senses, with vision providing global information about the scene and touch measuring local information during manipulation without suffering from occlusions. While prior work demonstrates the efficacy of tactile sensing for precise manip…

Cited by 33SourcePDFScholar
2023

Semantic Mechanical Search with Large Vision and Language Models

CoRL 2023poster

Moving objects to find a fully-occluded target object, known as mechanical search, is a challenging problem in robotics. As objects are often organized semantically, we conjecture that semantic information about object relationships can facilitate mechanical search and reduce search time. Large pret…

Cited by 11SourceScholar
2022

All You Need is LUV: Unsupervised Collection of Labeled Images Using UV-Fluorescent Markings

IROS 2022poster

Learning-based perception systems in robotics often requires large-scale image segmentation annotation. Current approaches rely on human labelers, which can be expensive, or simulation data, which can visually differ from real data. This paper proposes Labels from UltraViolet (LUV), a novel framewor…

Cited by 12SourceScholar
2022

Evo-NeRF: Evolving NeRF for Sequential Robot Grasping of Transparent Objects

CoRL 2022oral

Sequential robot grasping of transparent objects, where a robot removes objects one by one from a workspace, is important in many industrial and household scenarios. We propose Evolving NeRF (Evo-NeRF), leveraging recent speedups in NeRF training and further extending it to rapidly train the NeRF re…

Cited by 100SourceScholar
2022

Mechanical Search on Shelves using a Novel “Bluction” Tool

ICRA 2022poster

Shelves are common in homes, warehouses, and commercial settings due to their storage efficiency. However, this efficiency comes at the cost of reduced visibility and accessibility. When looking from a side (lateral) view of a shelf, most objects will be fully occluded, resulting in a constrained la…

Cited by 24SourceScholar
2022

Real2Sim2Real: Self-Supervised Learning of Physical Single-Step Dynamic Actions for Planar Robot Casting

ICRA 2022poster

This paper introduces the task of Planar Robot Casting (PRC): where one planar motion of a robot arm holding one end of a cable causes the other end to slide across the plane toward a desired target. PRC allows the cable to reach points beyond the robot workspace and has applications for cable manag…

Cited by 70SourceScholar
2021

Learning Seed Placements and Automation Policies for Polyculture Farming with Companion Plants

ICRA 2021poster

Polyculture farming is a sustainable farming technique based on synergistic interactions between differing plant types that make them more resistant to diseases and pests and better able to retain water. Reduced uniformity can reduce use of pesticides, fertilizer, and water, but is more labor intens…

Cited by 16SourcecodeScholar
2021

Mechanical Search on Shelves using Lateral Access X-RAY

IROS 2021poster

Finding an occluded object in a lateral access environment such as a shelf or cabinet is a problem that arises in many contexts such as warehouses, retail, healthcare, shipping, and homes. While this problem, known as mechanical search, is well-studied in overhead access environments, lateral access…

Cited by 32SourceScholar
2021

PREGAN: Pose Randomization and Estimation for Weakly Paired Image Style Translation

RA-L 2021

Utilizing the trained model under different conditions without data annotation is attractive for robot applications. Towards this goal, one class of methods is to translate the image style from another environment to the one on which models are trained. In this letter, we propose a weakly-paired set

Cited by 1SourcecodeScholar
2021

REDE: End-to-End Object 6D Pose Robust Estimation Using Differentiable Outliers Elimination

RA-L 2021

Object 6D pose estimation is a fundamental task in many applications. Conventional methods solve the task by detecting and matching the keypoints, then estimating the pose. Recent efforts bringing deep learning into the problem mainly overcome the vulnerability of conventional methods to environment

Cited by 41SourcecodeScholar
2021

Robots of the Lost Arc: Self-Supervised Learning to Dynamically Manipulate Fixed-Endpoint Cables

ICRA 2021poster

We explore how high-speed robot arm motions can dynamically manipulate ropes and cables to vault over obstacles, knock objects from pedestals, and weave between obstacles. In this paper, we propose a self-supervised learning framework that enables a UR5 robot to perform these three tasks. The framew…

Cited by 72SourceScholar
2019

Complex Stiffness Model of Physical Human-Robot Interaction: Implications for Control of Performance Augmentation Exoskeletons

IROS 2019poster

Human joint dynamic stiffness plays an important role in the stability of performance augmentation exoskeletons. In this paper, we consider a new frequency domain model of the human joint dynamics which features a complex value stiffness. This complex stiffness consists of a real stiffness and a hys…

Cited by 8SourceScholar