← Search

Zhixiong Zeng

9 accepted papers

2026

Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation

ICLR 2026poster

While reinforcement learning (RL) has proven highly effective for general reasoning in vision-language models, its application to tasks requiring deep understanding of information-rich images and structured output generation remains underexplored. Chart-to-code generation exemplifies this challenge,…

Cited by 0SourcecodeScholar
2026

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

ICLR 2026poster

Multimodal large language models are progressively advancing toward multimodal agents that can proactively execute tasks. Existing research on multimodal agents primarily targets either GUI or embodied scenarios, corresponding to interactions within 2D virtual world and 3D physical world, respective…

Cited by 0SourceScholar
2026

Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs) and is now being applied to Vision-Language Models (VLMs). However, vanilla RLVR for VLMs verifies only the final textual output, critically neglecting the foun

Cited by 0SourcecodeScholar
2026

Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR

CVPR 2026

Reading text from images or scanned documents via OCR models has been a longstanding focus of researchers. Intuitively, text reading is perceived as a straightforward perceptual task, and existing work primarily focuses on constructing enriched data engineering to enhance SFT capabilities. In this w

Cited by 0SourcecodeScholar
2026

SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start

ICLR 2026poster

Reinforcement learning (RL) with verifiable rewards has recently catalyzed a wave of “MLLM-r1” approaches that bring RL to vision language models. Most representative paradigms begin with a cold start, typically employing supervised fine-tuning (SFT), to initialize the policy before RL. However, SFT…

Cited by 0SourcecodeScholar
2026

TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution

ICML 2026poster

Effectively scaling GUI automation is essential for computer-use agents (CUAs); however, existing work primarily focuses on scaling GUI grounding rather than the more crucial GUI planning, which requires more sophisticated data collection. In reality, the exploration process of a CUA across apps/des…

Cited by 0SourceScholar
2025

Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy

NeurIPS 2025poster

In this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower information density and non-uniform spatial distribution. Accordingly, we present an entropy-informed decoding strategy that fac…

Cited by 0SourceScholar
2024

GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning

NeurIPS 2024poster

Large Language Models (LLMs) are increasingly used for various tasks with graph structures. Though LLMs can process graph information in a textual format, they overlook the rich vision modality, which is an intuitive way for humans to comprehend structural information and conduct general graph reaso…

2024

Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging

EMNLP 2024main

Supervised fine-tuning (SFT) is crucial for adapting Large Language Models (LLMs) to specific tasks. In this work, we demonstrate that the order of training data can lead to significant training imbalances, potentially resulting in performance degradation. Consequently, we propose to mitigate this i…

Cited by 1SourcePDFScholar