← Search

Zhiheng Wu

4 accepted papers

2026

M2GRPO: Mamba-Based Multi-Agent Group Relative Policy Optimization for Biomimetic Underwater Robots Pursuit

ICRA 2026poster

Traditional policy learning methods in cooperative pursuit face fundamental challenges in biomimetic underwater robots, where long-horizon decision making, partial observability, and inter-robot coordination require both expressiveness and stability. To address these issues, a novel framework called…

Cited by 0Scholar
2026

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection

CVPR 2026

Recent advances in Vision-Language Models (VLMs) have benefited from Reinforcement Learning (RL) for enhanced reasoning. However, existing methods still face critical limitations, including the lack of low-level visual information and effective visual feedback. To address these problems, this paper

Cited by 0SourceScholar
2023

PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction

AAAI 2023technical

Multi-task deep reinforcement learning (DRL) ambitiously aims to train a general agent that masters multiple tasks simultaneously. However, varying learning speeds of different tasks compounding with negative gradients interference makes policy learning inefficient. In this work, we propose PiCor, a…

2022

UC-OWOD: Unknown-Classified Open World Object Detection

ECCV 2022poster

"Open World Object Detection (OWOD) is a challenging computer vision problem that requires detecting unknown objects and gradually learning the identified unknown classes. However, it cannot distinguish unknown instances as multiple unknown classes. In this work, we propose a novel OWOD problem call…