← Search

Peizhou Cao

2 accepted papers

2026

G$^2$TAM: Geometry Grounded Track Anything Model

ICML 2026poster

Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explicit object appearance memory banks for instance tracking, yet they remain vulnera…

Cited by 0SourceScholar
2025

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

NeurIPS 2025poster

The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts. However, existing datasets typically suffer from limitations in data scale or diversity, sanitized layouts lacking small items, and severe object collis…

Cited by 0SourceScholar