2025
ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
NeurIPS 2025spotlight
Multi-modal large language models (MLLMs) are making rapid progress toward general-purpose embodied agents. However, existing MLLMs do not reliably capture fine-grained links between low-level visual features and high-level textual semantics, leading to weak grounding and inaccurate perception. To o…