SERL: Self-Examining Reinforcement Learning on Open-Domain
Weixuan Ou, Yanzhao Zheng, Shuoshuo Sun, Wei Zhang, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu
Abstract
Reinforcement Learning (RL) has been shown to improve the capabilities of large language models (LLMs). However, applying RL to open-domain tasks faces two key challenges: (1) the inherent subjectivity of these tasks prevents the verifiable rewards as required by Reinforcement Learning with Verifiable Rewards (RLVR); (2) Reinforcement Learning from Human Feedback (RLHF) relies on external reward mechanisms. To overcome these limitations, we propose Self-Examining Reinforcement Learning (SERL), a novel self-improving framework where the LLM serves as both Actor and Judge. SERL introduces two synergistic reward mechanisms without any external signals. On the one hand, to improve the Actor
BibTeX
@inproceedings{aaai2026_serlselfexaminin,
title = {SERL: Self-Examining Reinforcement Learning on Open-Domain},
author = {Weixuan Ou and Yanzhao Zheng and Shuoshuo Sun and Wei Zhang and Baohua Dong and Hangcheng Zhu and Ruohui Huang and Gang Yu and Pengwei Yan and Yifan Qiao},
booktitle = {AAAI 2026},
year = {2026}
}