← Search

Weixuan Ou

1 accepted papers

2026

SERL: Self-Examining Reinforcement Learning on Open-Domain

AAAI 2026technical

Reinforcement Learning (RL) has been shown to improve the capabilities of large language models (LLMs). However, applying RL to open-domain tasks faces two key challenges: (1) the inherent subjectivity of these tasks prevents the verifiable rewards as required by Reinforcement Learning with Verifiab

Cited by 0SourcePDFScholar