← Search

Haonan Chang

15 accepted papers

2025

Autoregressive Action Sequence Learning for Robotic Manipulation

RA-L 2025

Designing a universal policy architecture that performs well across diverse robots and task configurations remains a key challenge. In this work, we address this by representing robot actions as sequential data and generating actions through autoregressive sequence modeling. Existing autoregressive

Cited by 37SourcecodeScholar
2025

Failure Forecasting Boosts Robustness of Sim2Real Rhythmic Insertion Policies

IROS 2025

This paper addresses the challenges of Rhythmic Insertion Tasks (RIT), where a robot must repeatedly perform high-precision insertions, such as screwing a nut into a bolt with a wrench. The inherent difficulty of RIT lies in achieving millimeter-level accuracy and maintaining consistent performance

Cited by 0SourcecodeScholar
2025

UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models

ICRA 2025

Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified f

Cited by 10SourceScholar
2024

A3VLM: Actionable Articulation-Aware Vision Language Model

CoRL 2024poster

Vision Language Models (VLMs) for robotics have received significant attention in recent years. As a VLM can understand robot observations and perform complex visual reasoning, it is regarded as a potential universal solution for general robotics challenges such as manipulation and navigation. Howev…

Cited by 12SourcecodeScholar
2024

DAP: Diffusion-based Affordance Prediction for Multi-modality Storage

IROS 2024poster

Solving storage problems—where objects must be accurately placed into containers with precise orientations and positions—presents a distinct challenge that extends beyond traditional rearrangement tasks. These challenges are primarily due to the need for fine-grained 6D manipulation and the inherent…

Cited by 1SourcecodeScholar
2024

Insert-One: One-Shot Robust Visual-Force Servoing for Novel Object Insertion with 6-DoF Tracking

IROS 2024poster

Recent advancements in autonomous robotic assembly have shown promising results, especially in addressing the precision insertion challenge. However, achieving adaptability across diverse object categories and tasks often necessitates a learning phase that requires costly real-world data collection.…

Cited by 2SourceScholar
2024

LGMCTS: Language-Guided Monte-Carlo Tree Search for Executable Semantic Object Rearrangement

IROS 2024

We present LGMCTS, a framework that uniquely combines language guidance with geometrically informed sampling distributions to effectively rearrange objects according to geometric patterns dictated by natural language descriptions. LGMCTS uses Monte Carlo Tree Search (MCTS) to create feasible action

Cited by 17SourcecodeScholar
2024

Scaling Manipulation Learning with Visual Kinematic Chain Prediction

CoRL 2024poster

Learning general-purpose models from diverse datasets has achieved great success in machine learning. In robotics, however, existing methods in multi-task learning are typically constrained to a single robot and workspace, while recent work such as RT-X requires a non-trivial action normalization pr…

Cited by 1SourcecodeScholar
2023

Context-Aware Entity Grounding with Open-Vocabulary 3D Scene Graphs

CoRL 2023poster

We present an Open-Vocabulary 3D Scene Graph (OVSG), a formal framework for grounding a variety of entities, such as object instances, agents, and regions, with free-form text-based queries. Unlike conventional semantic-based object localization approaches, our system facilitates context-aware entit…

Cited by 29SourcecodeScholar
2023

Mono-STAR: Mono-Camera Scene-Level Tracking and Reconstruction

ICRA 2023poster

We present Mono-STAR, the first real-time 3D reconstruction system that simultaneously supports semantic fusion, fast motion tracking, non-rigid object deformation, and topological change under a unified framework. The proposed system solves a new optimization problem incorporating optical-flow-base…

Cited by 6SourcecodeScholar
2023

OVIR-3D: Open-Vocabulary 3D Instance Retrieval Without Training on 3D Data

CoRL 2023poster

This work presents OVIR-3D, a straightforward yet effective method for open-vocabulary 3D object instance retrieval without using any 3D data for training. Given a language query, the proposed method is able to return a ranked set of 3D object instance segments based on the feature similarity of the…

Cited by 61SourcecodeScholar
2020

GeoFusion: Geometric Consistency Informed Scene Estimation in Dense Clutter

RA-L 2020

We propose GeoFusion, a SLAM-based scene estimation method for building an object-level semantic map in dense clutter. In dense clutter, objects are often in close contact and severe occlusions, which brings more false detections and noisy pose estimates from existing perception methods. To solve th

Cited by 10SourceScholar
2019

GlassLoc: Plenoptic Grasp Pose Detection in Transparent Clutter

IROS 2019poster

Transparent objects are prevalent across many environments of interest for dexterous robotic manipulation. Such transparent material leads to considerable uncertainty for robot perception and manipulation, and remains an open challenge for robotics. This problem is exacerbated when multiple transpar…

Cited by 28SourceScholar