← Search

Yinghong Liao

6 accepted papers

2026

PhyScene3D: Physically Consistent 3D Interactive Tabletop Scene Generation

ICML 2026poster

Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense object hierarchies and irregular affordances. Existing methods, ranging from decoupled symbolic solvers to end-to-end regress…

Cited by 0SourceScholar
2025

Empowering Large Language Models with 3D Situation Awareness

CVPR 2025poster

Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions…

Cited by 0SourcePDFScholar
2023

BEV@DC: Bird's-Eye View Assisted Training for Depth Completion

CVPR 2023poster

Depth completion plays a crucial role in autonomous driving, in which cameras and LiDARs are two complementary sensors. Recent approaches attempt to exploit spatial geometric constraints hidden in LiDARs to enhance image-guided depth completion. However, only low efficiency and poor generalization c…

Cited by 30SourcePDFScholar
2023

Geometry-Aware Network for Domain Adaptive Semantic Segmentation

AAAI 2023technical

Measuring and alleviating the discrepancies between the synthetic (source) and real scene (target) data is the core issue for domain adaptive semantic segmentation. Though recent works have introduced depth information in the source domain to reinforce the geometric and semantic knowledge transfer,…

Cited by 6SourcePDFScholar
2022

X-Trans2Cap: Cross-Modal Knowledge Transfer Using Transformer for 3D Dense Captioning

CVPR 2022poster

3D dense captioning aims to describe individual objects by natural language in 3D scenes, where 3D scenes are usually represented as RGB-D scans or point clouds. However, only exploiting single modal information, e.g., point cloud, previous approaches fail to produce faithful descriptions. Though ag…

Cited by 96PDFcodeScholar
2021

InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds Through Instance Multi-Level Contextual Referring

ICCV 2021poster

Compared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer, to achieve a superior 3D visual grounding through the grounding-by-matching strategy. In practice, our…

Cited by 147PDFcodeScholar