AffIn-Space: Learning Affine-Invariant Representations for 3D Spatial Understanding with MLLMs
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general visual understanding, they suffer from a fundamental geometric fragility: standard visual representations often degrade rapidly under changes in viewpoint and viewing distance. Our analysis identifies that ex…