← Search

Yuxi Xiao

9 accepted papers

2026

SpatialTree: How Spatial Intelligence Branches Out in MLLMs

CVPR 2026

Cognitive science suggests that spatial ability develops progressively--from perception to reasoning and interaction. Yet in multimodal LLMs (MLLMs), this hierarchy remains poorly understood, as most studies focus on a narrow set of tasks. We introduce SpatialTree, a cognitive-science-inspired hiera

Cited by 0SourcecodeScholar
2026

Trace Anything: Representing Any Video in 4D via Trajectory Fields

ICLR 2026poster

Building 4D video representations to model underlying spacetime constitutes a crucial step toward understanding dynamic scenes, yet there is no consensus on the paradigm: current approaches resort to additional estimators such as depth, flow, or tracking, or to heavy per-scene optimization, making t…

Cited by 0SourcecodeScholar
2025

ERNet: Efficient Non-Rigid Registration Network for Point Sequences

ICCV 2025poster

Registering an object shape to a sequence of point clouds undergoing non-rigid deformation is a long-standing challenge. The key difficulties stem from two factors: (i) the presence of local minima due to the non-convexity of registration objectives, especially under noisy or partial inputs, which h…

Cited by 0SourcePDFScholar
2025

SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera Motion

ICCV 2025poster

We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-…

Cited by 0SourcePDFScholar
2024

CoDeF: Content Deformation Fields for Temporally Consistent Video Processing

CVPR 2024highlight

We present the content deformation field (CoDeF) as a new type of video representation which consists of a canonical content field aggregating the static contents in the entire video and a temporal deformation field recording the transformations from the canonical image (i.e. rendered from the canon…

2024

NEAT: Distilling 3D Wireframes from Neural Attraction Fields

CVPR 2024poster

This paper studies the problem of structured 3D recon- struction using wireframes that consist of line segments and junctions focusing on the computation of structured boundary geometries of scenes. Instead of leveraging matching-based solutions from 2D wireframes (or line segments) for 3D wireframe…

2024

SpatialTracker: Tracking Any 2D Pixels in 3D Space

CVPR 2024highlight

Recovering dense and long-range pixel motion in videos is a challenging problem. Part of the difficulty arises from the 3D-to-2D projection process leading to occlusions and discontinuities in the 2D motion domain. While 2D motion can be intricate we posit that the underlying 3D motion can often be…

2023

Level-S$^2$fM: Structure From Motion on Neural Level Set of Implicit Surfaces

CVPR 2023poster

This paper presents a neural incremental Structure-from-Motion (SfM) approach, Level-S2fM, which estimates the camera poses and scene geometry from a set of uncalibrated images by learning coordinate MLPs for the implicit surfaces and the radiance fields from the established keypoint correspondences…

2022

DeepMLE: A Robust Deep Maximum Likelihood Estimator for Two-view Structure from Motion

IROS 2022poster

Two-view structure from motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM (vSLAM). Many existing end-to-end learning-based methods usually formulate it as a brute regression problem. However, the inadequate utilization of traditional geometry model makes the model not robust in un…

Cited by 10SourceScholar