2025
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
NeurIPS 2025poster
Visual reasoning in multimodal large language models (MLLMs) has primarily been studied in passive, static settings, limiting their effectiveness in real-world physical environments where an embodied agent must contend with incomplete information due to occlusion or a limited field of view. Humans,…