← Search

Yuxi Zheng

2 accepted papers

2026

Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning

CVPR 2026

Multimodal Large Language Models (MLLMs) have made great progress in video understanding tasks. However, when it comes to understanding complex or lengthy videos, MLLMs tend to overlook details or produce hallucinations. To alleviate these issues, recent work has attempted to leverage reinforcement

Cited by 0SourcecodeScholar
2025

Robot Learning from Any Images

CoRL 2025poster

We introduce RoLA, a framework that transforms any in‑the‑wild image into an interactive, physics‑enabled robotic environment. Unlike previous methods, RoLA operates directly on a single image without requiring additional hardware or digital assets. Our framework democratizes robotic data generatio…

Cited by 0SourcecodeScholar