← Search

Jaehwi Song

1 accepted papers

2026

Learning Multi-View Spatial Reasoning from Cross-View Relations

CVPR 2026

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manipulate objects across different viewpoints. In this work, we introduce Cross-View

Cited by 0SourceScholar