← Search

Qinghao Wang

2 accepted papers

2026

RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-Training

ICLR 2026poster

Reinforcement learning with verifiable reward has recently emerged as a central paradigm for post-training large language models (LLMs); however, prevailing mean-based methods, such as Group Relative Policy Optimization (GRPO), suffer from entropy collapse and limited reasoning gains. We argue that…

Cited by 0SourcecodeScholar
2020

An Electrocommunication System Using FSK Modulation and Deep Learning Based Demodulation for Underwater Robots

IROS 2020poster

Underwater communication is extremely challenging for small underwater robots which typically have stringent power and size constraints. In our previous work, we developed an artificial electrocommunication system which could be an alternative for the communication of small underwater robots. This p…

Cited by 7SourceScholar