← Search

Wentian Zhao

9 accepted papers

2026

HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation

ICLR 2026poster

Agentic Retrieval-Augmented Generation (RAG) is a powerful technique for incorporating external information that Large Language Models (LLMs) lack, enabling better problem solving and question answering. However, suboptimal search behaviors exist widely, such as over-search (retrieving information a…

Cited by 0SourceScholar
2026

Seeing is Solving: Unlocking Efficient Multimodal RL via View Alignment

ICML 2026poster

Although Reinforcement Learning Fine-Tuning (RLFT) applied to Vision-Language Models (VLMs) substantially enhances multimodal reasoning capabilities, their prohibitive training cost limits broad adoption. Surprisingly, most existing methods simply port Large Language Model (LLM) RLFT techniques to V…

Cited by 0SourceScholar
2026

Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play

ICLR 2026poster

Although reinforcement learning (RL) can effectively enhance the reasoning capabilities of vision–language models (VLMs), current methods remain heavily dependent on labor-intensive datasets that require extensive manual construction and verification, leading to extremely high training costs and con…

Cited by 0SourcecodeScholar
2025

Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference

NeurIPS 2025oral

Large Language Models (LLMs) are now integral across various domains and have demonstrated impressive performance. Progress, however, rests on the premise that benchmark scores are both accurate and reproducible. We demonstrate that the reproducibility of LLM performance is fragile: changing system…

Cited by 0SourcecodeScholar
2024

DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision

CVPR 2024poster

We have witnessed significant progress in deep learning-based 3D vision ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However existing scene-level datasets for deep learning-based 3D vision limited to either synthetic enviro…

Cited by 85SourcePDFScholar
2024

Relational Distant Supervision for Image Captioning without Image-Text Pairs

AAAI 2024technical

Unsupervised image captioning aims to generate descriptions of images without relying on any image-sentence pairs for training. Most existing works use detected visual objects or concepts as bridge to connect images and texts. Considering that the relationship between objects carries more informatio…

Cited by 5SourcePDFScholar
2019

Joint Syntax Representation Learning and Visual Cue Translation for Video Captioning

ICCV 2019poster

Video captioning is a challenging task that involves not only visual perception but also syntax representation learning. Recent progress in video captioning has been achieved through visual perception, but syntax representation learning is still under-explored. We propose a novel video captioning ap…

Cited by 112PDFScholar