← Search

Yufeng Zhong

5 accepted papers

2026

Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation

ICLR 2026poster

While reinforcement learning (RL) has proven highly effective for general reasoning in vision-language models, its application to tasks requiring deep understanding of information-rich images and structured output generation remains underexplored. Chart-to-code generation exemplifies this challenge,…

Cited by 0SourcecodeScholar
2026

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

ICLR 2026poster

Multimodal large language models are progressively advancing toward multimodal agents that can proactively execute tasks. Existing research on multimodal agents primarily targets either GUI or embodied scenarios, corresponding to interactions within 2D virtual world and 3D physical world, respective…

Cited by 0SourceScholar
2026

Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR

CVPR 2026

Reading text from images or scanned documents via OCR models has been a longstanding focus of researchers. Intuitively, text reading is perceived as a straightforward perceptual task, and existing work primarily focuses on constructing enriched data engineering to enhance SFT capabilities. In this w

Cited by 0SourcecodeScholar
2025

RoboTrom-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction

ICCV 2025poster

In language-guided visual navigation, agents locate target objects in unseen environments using natural language instructions. For reliable navigation in unfamiliar scenes, agents should possess strong perception, planning, and prediction capabilities. Additionally, when agents revisit previously ex…

Cited by 0SourcePDFScholar
2025

RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

ICCV 2025poster

Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial interactions and managing data collection costs. To address these issues, we propose the multimodal robotic manipulation…