← Search

Haoang Li

43 accepted papers

2026

Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics

ICML 2026poster

Video action models are a promising foundation for Vision–Language–Action (VLA) because they can learn rich visual dynamics directly from video. However, likelihood-oriented training of diffusion predictors emphasizes globally plausible futures and does not guarantee precision-critical visual dynami…

Cited by 0SourceScholar
2026

Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using A Feed-Forward 3D Model

RSS 2026poster

Fast and reliable initialization is critical for monocular visual–inertial navigation systems (VINS), as it establishes the starting conditions for subsequent state estimation. Despite steady progress, most existing methods heavily rely on visual feature correspondences and require 3-4 seconds of se…

Cited by 0SourceScholar
2026

Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning

ICRA 2026poster

Whole-body manipulation (WBM) for humanoid robots presents a promising approach for executing embracing tasks involving bulky objects, where traditional grasping relying on end-effectors only remains limited in such scenarios due to inherent stability and payload constraints. This paper introduces a…

2026

Neural Predictor-Corrector: Solving Homotopy Problems with Reinforcement Learning

ICLR 2026poster

The Homotopy paradigm, a general principle for solving challenging problems, appears across diverse domains such as robust optimization, global optimization, polynomial root-finding, and sampling. Practical solvers for these problems typically follow a predictor-corrector (PC) structure, but rely on…

Cited by 0SourceScholar
2026

Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation

CVPR 2026

Generating realistic hand-object interactions (HOI) videos is a significant challenge due to the difficulty of modeling physical constraints (e.g., contact and occlusion between hands and manipulated objects). Current methods utilize HOI representation as an auxiliary generative objective to guide v

Cited by 0SourceScholar
2026

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

AAAI 2026technical

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispe

Cited by 0SourcePDFScholar
2026

Rethinking the Practicality of Vision-Language-Action Model: A Comprehensive Benchmark and an Improved Baseline

ICRA 2026poster

Vision-Language-Action (VLA) models have emerged as a generalist robotic agent. However, existing VLAs are hindered by excessive parameter scales, prohibitive pre-training requirements, and limited applicability to diverse embodiments. To improve the practicality of VLAs, we propose a comprehensive …

2026

SVP: Improving Vision-Language-Action Models with Dual Stochastic Visual Prompting

ICRA 2026poster

Vision-Language-Action (VLA) models, such as OpenVLA, hold the promise of generalist robots, yet their performance is often impaired by distracted attention, which we identify as a manifestation of shortcut learning. We posit that the solution lies not in architectural modifications, but in a new tr…

Cited by 0Scholar
2026

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

ICLR 2026poster

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data, which lack accurate spatial awareness and hinder their abili…

Cited by 0SourcecodeScholar
2026

Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Diffusion Diffusion Process

ICLR 2026poster

Vision-language-action (VLA) models aim to understand natural language instructions and visual observations and execute corresponding actions as an embodied agent. Recent advancements have integrated future images into the understanding-action loop, enabling foresight-driven policies that reduce abs…

Cited by 0SourcecodeScholar
2025

Convex Relaxation for Robust Vanishing Point Estimation in Manhattan World

CVPR 2025award

Determining the vanishing points (VPs) in a Manhattan world, as a fundamental task in many 3D vision applications, consists of jointly inferring the line-VP association and locating each VP. Existing methods are, however, either sub-optimal solvers or pursuing global optimality at a significant cost…

2025

G2-SDF: Geometry-Guided Neural Signed Distance Fields for Scalable and Detailed Reconstruction

RA-L 2025

Effcient reconstruction methods, particularly capable of providing detailed information on obstacle distances across diverse environments, are crucial for effective robot motion planning. In this context, neural Signed Distance Fields (SDFs) offer a powerful solution by learning implicit representat

Cited by 1SourceScholar
2025

Interactive Navigation for Legged Manipulators with Learned Arm-Pushing Controller

IROS 2025

Interactive navigation is crucial in scenarios where proactively interacting with objects can yield shorter paths, thus significantly improving traversal efficiency. Existing methods primarily focus on using the robot body to relocate obstacles during navigation. However, they prove ineffective in n

Cited by 5SourcecodeScholar
2025

L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR Model

IROS 2025

Semantic Scene Completion (SSC) constitutes a pivotal element in autonomous driving perception systems, tasked with inferring the 3D semantic occupancy of a scene from sensory data. To improve accuracy, prior research has implemented various computationally demanding and memory-intensive 3D operatio

Cited by 3SourcecodeScholar
2025

PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding

IROS 2025

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in

Cited by 60SourceScholar
2025

RMG: Real-Time Expressive Motion Generation with Self-collision Avoidance for 6-DOF Companion Robotic Arms

IROS 2025

The six-degree-of-freedom (6-DOF) robotic arm has gained widespread application in human-coexisting environments. While previous research has predominantly focused on functional motion generation, the critical aspect of expressive motion in human-robot interaction remains largely unexplored. This pa

Cited by 0SourceScholar
2025

RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation

IROS 2025

This paper introduces RoboDexVLM, an innovative framework for robot task planning and grasp detection tailored for a collaborative manipulator equipped with a dexterous hand. Previous methods focus on simplified and limited manipulation tasks, which often neglect the complexities associated with gra

Cited by 15SourcecodeScholar
2025

STG-Avatar: Animatable Human Avatars via Spacetime Gaussian

IROS 2025

Realistic animatable human avatars from monocular videos are crucial for advancing human-robot interaction and enhancing immersive virtual experiences. While recent research on 3DGS-based human avatars has made progress, it still struggles with accurately representing detailed features of non-rigid

Cited by 6SourcecodeScholar
2025

San Francisco World: Leveraging Structural Regularities of Slope for 3-DoF Visual Compass

RA-L 2025

We propose the San Francisco world (SFW) model, a novel structural model inspired by San Francisco's hilly terrain, enabling 3D inter-floor navigation in urban areas rather than being limited to 2D intra-floor navigation of various robotics platforms. Our SFW consists of a single vertical dominant d

Cited by 4SourcecodeScholar
2025

SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments

IROS 2025

Unmanned Aerial Vehicles (UAVs) have emerged as versatile tools across various sectors, driven by their mobility and adaptability. This paper introduces SkyVLN, a novel framework integrating vision-and-language navigation (VLN) with Nonlinear Model Predictive Control (NMPC) to enhance UAV autonomy i

Cited by 12SourceScholar
2024

GlobalPointer: Large-Scale Plane Adjustment with Bi-Convex Relaxation

ECCV 2024poster

"Plane adjustment (PA) is crucial for many 3D applications, involving simultaneous pose estimation and plane recovery. Despite recent advancements, it remains a challenging problem in the realm of multi-view point cloud registration. Current state-of-the-art methods can achieve globally optimal conv…

2024

Physically-Based Photometric Bundle Adjustment in Non-Lambertian Environments

IROS 2024poster

Photometric bundle adjustment (PBA) is widely used in estimating the camera pose and 3D geometry by assuming a Lambertian world. However, the assumption of photometric consistency is often violated since the non-diffuse reflection is common in real-world environments. The photometric inconsistency s…

Cited by 0SourceScholar
2024

Structured-NeRF: Hierarchical Scene Graph with Neural Representation

ECCV 2024poster

"We present Structured Neural Radiance Field (Structured-NeRF) for indoor scene representaion based on a novel hierarchical scene graph structure to organize the neural radiance field. Existing object-centric methods focus only on the inherent characteristics of objects, while overlooking the semant…

Cited by 2SourcePDFScholar
2023

DDIT: Semantic Scene Completion via Deformable Deep Implicit Templates

ICCV 2023poster

Scene reconstructions are often incomplete due to occlusions and limited viewpoints. There have been efforts to use semantic information for scene completion. However, the completed shapes may be rough and imprecise since respective methods rely on 3D convolution and/or lack effective shape constrai…

Cited by 10PDFScholar
2023

Learning Accurate 3D Shape Based on Stereo Polarimetric Imaging

CVPR 2023poster

Shape from Polarization (SfP) aims to recover surface normal using the polarization cues of light. The accuracy of existing SfP methods is affected by two main problems. First, the ambiguity of polarization cues partially results in false normal estimation. Second, the widely-used assumption about o…

Cited by 12SourcePDFScholar
2022

Quasi-Globally Optimal and Real-Time Visual Compass in Manhattan Structured Environments

RA-L 2022

We present a drift-free visual compass for estimating the three degrees of freedom (DoF) rotational motion of a camera by recognizing structural regularities in a Manhattan world (MW), which posits that the major structures conform to three orthogonal principal directions. Existing Manhattan frame e

Cited by 12SourcecodeScholar
2021

Learning Icosahedral Spherical Probability Map Based on Bingham Mixture Model for Vanishing Point Estimation

ICCV 2021poster

Existing vanishing point (VP) estimation methods rely on pre-extracted image lines and/or prior knowledge of the number of VPs. However, in practice, this information may be insufficient or unavailable. To solve this problem, we propose a network that treats a perspective image as input and predicts…

Cited by 9PDFScholar
2021

Learning To Identify Correct 2D-2D Line Correspondences on Sphere

CVPR 2021poster

Given a set of putative 2D-2D line correspondences, we aim to identify correct matches. Existing methods exploit the geometric constraints. They are only applicable to structured scenes with orthogonality, parallelism and coplanarity. In contrast, we propose the first approach suitable for both stru…

Cited by 4PDFScholar
2021

Task-Space Trajectory Tracking Control for Coordinated Manipulation Using Sampled Coupling Data

RA-L 2021

This letter studies the task-space synchronization control problem of networked manipulators which are commanded to track the desired task-space trajectories to achieve the coordinated manipulation transportation tasks in industrial and logistic applications. To guarantee the practical applicability

Cited by 8SourceScholar
2020

A Synchronization Approach for Achieving Cooperative Adaptive Cruise Control Based Non-Stop Intersection Passing

ICRA 2020poster

Cooperative adaptive cruise control (CACC) of intelligent vehicles contributes to improving cruise control performance, reducing traffic congestion, saving energy and increasing traffic flow capacity. In this paper, we resolve the CACC problem from the viewpoint of synchronization control, our main…

Cited by 7SourceScholar
2020

CUHK-AHU Dataset: Promoting Practical Self-Driving Applications in the Complex Airport Logistics, Hill and Urban Environments

IROS 2020poster

This paper presents a novel dataset targeting three types of challenging environments for autonomous driving, i.e., the industrial logistics environment, the undulating hill environment and the mixed complex urban environment. To the best of the author’s knowledge, similar dataset has not been publi…

Cited by 5SourceScholar
2020

End-to-End 3D Point Cloud Learning for Registration Task Using Virtual Correspondences

IROS 2020poster

3D Point cloud registration is still a very challenging topic due to the difficulty in finding the rigid transformation between two point clouds with partial correspondences, and it's even harder in the absence of any initial estimation information. In this paper, we present an end-to-end deep-learn…

Cited by 26SourcecodeScholar
2020

Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World

ECCV 2020poster

Atlanta world holds for the scenes composed of a vertical dominant direction and several horizontal dominant directions. Vanishing point (VP) is the intersection of the image lines projected from parallel 3D lines. In Atlanta world, given a set of image lines, we aim to cluster them by the unknown-b…

Cited by 17SourcePDFScholar
2020

Robust and Efficient Estimation of Absolute Camera Pose for Monocular Visual Odometry

ICRA 2020poster

Given a set of 3D-to-2D point correspondences corrupted by outliers, we aim to robustly estimate the absolute camera pose. Existing methods robust to outliers either fail to guarantee high robustness and efficiency simultaneously, or require an appropriate initial pose and thus lack generality. In c…

Cited by 5SourceScholar
2019

A Hierarchical Framework for Coordinating Large-Scale Robot Networks

ICRA 2019poster

In this paper, we study the cooperative path planning and motion coordination problems of the multi-robot system with large number of robots, aiming for practical applications in robotic warehouses and automated transportation systems. Particularly, we solve the life-long planning problem and guaran…

Cited by 15SourceScholar
2019

LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment Analysis

ICCV 2019poster

Point cloud based place recognition is still an open issue due to the difficulty in extracting local features from the raw 3D point cloud and generating the global descriptor, and it's even harder in the large-scale dynamic environments. In this paper, we develop a novel deep neural network, named L…

Cited by 345PDFScholar
2019

Leveraging Structural Regularity of Atlanta World for Monocular SLAM

ICRA 2019poster

A wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocul…

Cited by 48SourceScholar
2019

Line-based Absolute and Relative Camera Pose Estimation in Structured Environments

IROS 2019poster

3D lines in structured environments encode particular regularity like parallelism and orthogonality. We leverage this structural regularity to estimate the absolute and relative camera poses. We decouple the rotation and translation, and propose a novel rotation estimation method. We decompose the a…

Cited by 22SourceScholar
2019

Quasi-Globally Optimal and Efficient Vanishing Point Estimation in Manhattan World

ICCV 2019oral

The image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim at clustering them by three unknown-but-sought VPs. The…

Cited by 37PDFScholar
2018

A Monocular SLAM System Leveraging Structural Regularity in Manhattan World

ICRA 2018poster

The structural features in Manhattan world encode useful geometric information of parallelism, orthogonality and/or coplanarity in the scene. By fully exploiting these structural features, we propose a novel monocular SLAM system which provides accurate estimation of camera poses and 3D map. The for…

Cited by 72SourceScholar
2018

Robust Camera Pose Estimation via Consensus on Ray Bundle and Vector Field

IROS 2018poster

Estimating the camera pose requires point correspondences. However, in practice, correspondences are inevitably corrupted by outliers, which affects the pose estimation. We propose a general and accurate outlier removal strategy for robust camera pose estimation. The proposed strategy can detect out…

Cited by 7SourceScholar
2017

Combining points and lines for camera pose estimation and optimization in monocular visual odometry

IROS 2017poster

In this paper, we propose a unified model for camera pose estimation and a novel strategy for pose optimization by combining points and lines in monocular visual odometry. Our proposed unified model treats point and line features equivalently, which is applicable for all the minimal cases requiring…

Cited by 23SourceScholar