← Search

Sicheng Xu

13 accepted papers

2026

HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models

CVPR 2026

Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatial reasoning. In this paper, we propose a principled hierarchical framework that

Cited by 0SourcecodeScholar
2026

Native and Compact Structured Latents for 3D Generation

CVPR 2026

Recent advancements in 3D generative modeling have significantly improved the generation realism, yet the field is still hampered by existing representations, which struggle to capture assets with complex topologies and detailed appearance. This paper present an approach for learning a structured la

Cited by 0SourcecodeScholar
2026

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

CVPR 2026

Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a framework for streamable talking portrait video generation conditioned on speech audio and reference images. Designed m

Cited by 0SourceScholar
2026

Scalable Vision-Language-Action Model Pretraining for Robotic Dexterous Manipulation with Real-Life Human Activity Videos

ICRA 2026poster

This paper presents an approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot end-effector, we show that "in-the-wild" egocentric human videos wit…

Cited by 0Scholar
2026

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

ICLR 2026poster

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet, most evaluations of VLMs focus on single-view settings, leaving their ability t…

Cited by 0SourcecodeScholar
2025

Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

ICCV 2025poster

In this paper, we present a novel framework for video-to-4D generation that creates high-quality dynamic 3D content from single video inputs. Direct 4D diffusion modeling is extremely challenging due to costly data construction and the high-dimensional nature of jointly representing 3D shape, appear…

2025

MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details

NeurIPS 2025poster

We propose MoGe-2, an advanced open-domain geometry estimation model that recovers a metric-scale 3D point map of a scene from a single image. Our method builds upon the recent monocular geometry estimation approach, MoGe, which predicts affine-invariant point maps with unknown scales. We explore ef…

Cited by 0SourceScholar
2025

MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

CVPR 2025poster

We present MoGe, a powerful model for recovering 3D geometry from monocular open-domain images. Given a single image, our model directly predicts a 3D point map of the captured scene with an affine-invariant representation, which is agnostic to true global scale and shift. This new representation pr…

Cited by 25SourcePDFScholar
2025

Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning

ICRA 2025

Robotic manipulation in cluttered environments presents significant challenges, particularly when the clutter includes thin, deformable objects like cables, which complicate perception and decision-making processes. In the context of datacenters, the automation of networking tasks often involves the

Cited by 0SourceScholar
2025

Structured 3D Latents for Scalable and Versatile 3D Generation

CVPR 2025highlight

We introduce a novel 3D generation method for versatile and high-quality 3D asset creation.The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D Gaussians, and meshes. This is achieved by integrating a spar…

2025

VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image

NeurIPS 2025poster

We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human faces, and reconstructing an intricate 3D head avatar from a single portrait image. To accurately model expression deta…

Cited by 0SourceScholar
2024

VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

NeurIPS 2024oral

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but als…

Cited by 92SourcePDFScholar