← Search

Chenchen Liu

5 accepted papers

2026

From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors

ICLR 2026poster

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, o…

Cited by 0SourcecodeScholar
2024

3D Affordance Keypoint Detection for Robotic Manipulation

IROS 2024poster

This paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts’ functionality. The proposed approach provides direct information about what the potential use of objects is, as well as guidance on where and ho…

Cited by 0SourceScholar
2021

Learning 3-D Human Pose Estimation from Catadioptric Videos

IJCAI 2021poster

3-D human pose estimation is a crucial step for understanding human actions. However, reliably capturing precise 3-D position of human joints is non-trivial and tedious. Current models often suffer from the scarcity of high-quality 3-D annotated training data. In this work, we explore a novel way of…

Cited by 4SourcePDFScholar
2020

Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global Context

CVPR 2020poster

Video visual relation detection (VidVRD) aims to describe all interacting objects in a video. Different from relationships in static images, videos contain an addition temporal channel. A majority of existing works divide a video into short segments, predict relationships in each segment, and merge…

Cited by 92PDFScholar