2025
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
ICCV 2025poster
We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human states and behaviors, reason about its surrounding environment, and answer human-rela…