← Search

Haowen Gao

2 accepted papers

2026

DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage

ICLR 2026poster

Reinforcement learning (RL) with group relative policy optimization (GRPO) has become a widely adopted approach for enhancing the reasoning capabilities of multimodal large language models (MLLMs). While GRPO enables long-chain reasoning without a traditional critic model, it often suffers from spar…

Cited by 0SourcecodeScholar
2024

Differentiable Space Carving for 3D Reconstruction Using Imaging Sonar

RA-L 2024

Effective 3D reconstruction utilizing imaging sonars is vital for underwater robots, particularly in turbid water conditions. The absence of elevation angles in acoustic echo measurements significantly slows down the Neural Radiance Field (NeRF) method. This is attributed to the differentiable rende

Cited by 12SourceScholar