2026
Gold Points Sniper: Self-Guided Visual Reasoning in VLM for Fine-Grained Action Understanding
ICRA 2026poster
Robots operating in everyday environments must understand fine-grained human actions, intentions, and contextual cues from broad views where people occupy only small regions, a capability unmet by current systems. While open-vocabulary action recognition methods remain limited to assigning predefine…