← Search

Zhuoxu Huang

2 accepted papers

2026

Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning

ICLR 2026poster

Reinforcement Learning with verifiable rewards (RLVR) has emerged as a primary learning paradigm for enhancing the reasoning capabilities of multi-modal large language models (MLLMs). However, during RL training, the enormous state space of MLLM and sparse rewards often leads to entropy collapse, po…

Cited by 0SourcecodeScholar
2024

MF-MOS: A Motion-Focused Model for Moving Object Segmentation

ICRA 2024poster

Moving object segmentation (MOS) provides a reliable solution for detecting traffic participants and thus is of great interest in the autonomous driving field. Dynamic capture is always critical in the MOS problem. Previous methods capture motion features from the range images directly. Differently,…

Cited by 19SourcecodeScholar