← Search

Ruohui Huang

2 accepted papers

2026

SERL: Self-Examining Reinforcement Learning on Open-Domain

AAAI 2026technical

Reinforcement Learning (RL) has been shown to improve the capabilities of large language models (LLMs). However, applying RL to open-domain tasks faces two key challenges: (1) the inherent subjectivity of these tasks prevents the verifiable rewards as required by Reinforcement Learning with Verifiab

Cited by 0SourcePDFScholar
2025

Reason from Future: Reverse Thought Chain Enhances LLM Reasoning

ACL 2025finding

It has been demonstrated that carefully designed reasoning paradigms, like Chain-of-Thought(CoT) and Tree-of-Thought(ToT), can enhance the reasoning capabilities of small language models by detailed thinking and extensive thought searching, unbounded branching factors in the searching space create p…

Cited by 0SourcePDFScholar