← Search

Ruoxuan Zhang

2 accepted papers

2026

Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention

CVPR 2026

Vision-Language-Action (VLA) models have enabled notable progress in general-purpose robotic manipulation, yet their learned policies often exhibit variable execution quality. We attribute this variability to the mixed-quality nature of human demonstrations, where the implicit principles that govern

Cited by 0SourceScholar
2026

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

CVPR 2026

Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision-making, and existing benchmarks focus solely on human mental states while ignoring the agent's own perspective, hinderi

Cited by 0SourcecodeScholar