← Search

Lichao Huang

12 accepted papers

2026

SEM: Enhancing Spatial Understanding for Robust Robot Manipulation

ICRA 2026poster

A key challenge in robot manipulation lies in developing policy models with consistent spatial understanding—the ability to reason about 3D geometry, object relations, and robot state. Existing mainstream models take 2D images as input, without performing explicit 3D modeling, and thus lack spatial …

2025

BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence

CVPR 2025poster

In embodied intelligence systems, a key component is 3D perception algorithm, which enables agents to understand their surrounding environments. Previous algorithms primarily rely on point cloud, which, despite offering precise geometric information, still constrain perception performance due to inh…

2025

DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data

NeurIPS 2025poster

We present **DIPO**, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state. Compared to the single-image approach, our dual-image input imposes only a modest overhead for da…

Cited by 0SourcecodeScholar
2025

Generating Multimodal Driving Scenes via Next-Scene Prediction

CVPR 2025poster

Generative models in Autonomous Driving (AD) enable diverse scenario creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multi…

2024

EDA: Evolving and Distinct Anchors for Multimodal Motion Prediction

AAAI 2024technical

Motion prediction is a crucial task in autonomous driving, and one of its major challenges lands in the multimodality of future behaviors. Many successful works have utilized mixture models which require identification of positive mixture components, and correspondingly fall into two main lines: pre…

2024

WidthFormer: Toward Efficient Transformer-based BEV View Transformation

IROS 2024poster

We present WidthFormer, a novel transformer-based module to compute Bird’s-Eye-View (BEV) representations from multi-view cameras for real-time autonomous-driving applications. WidthFormer is computationally efficient, robust and does not require any special engineering effort to deploy. We first in…

Cited by 3SourcecodeScholar
2020

Image Super-Resolution With Cross-Scale Non-Local Attention and Exhaustive Self-Exemplars Mining

CVPR 2020poster

Deep convolution-based single image super-resolution (SISR) networks embrace the benefits of learning from large-scale external image resources for local recovery, yet most existing works have ignored the long-range feature-wise similarities in natural images. Some recent works have successfully lev…

Cited by 478PDFcodeScholar
2019

CCNet: Criss-Cross Attention for Semantic Segmentation

ICCV 2019poster

Full-image dependencies provide useful contextual information to benefit visual understanding problems. In this work, we propose a Criss-Cross Network (CCNet) for obtaining such contextual information in a more effective and efficient way. Concretely, for each pixel, a novel criss-cross attention mo…

Cited by 3729PDFcodeScholar
2017

Parse geometry from a line: Monocular depth estimation with partial laser observation

ICRA 2017poster

Many standard robotic platforms are equipped with at least a fixed 2D laser range finder and a monocular camera. Although those platforms do not have sensors for 3D depth sensing capability, knowledge of depth is an essential part in many robotics activities. Therefore, recently, there is an increas…

Cited by 132SourceScholar