← Search

Linglong Li

1 accepted papers

2026

Video Spatial Reasoning with Object-Centric 3D Rollout

AAAI 2026technical

Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoning—the ability to comprehend object locations, orientations, and inter-object relationships in dynamic 3D scenes—remains

Cited by 0SourcePDFScholar