← Search

Alfred O. Hero

1 accepted papers

2026

CyclicReflex: Improving Reasoning Models via Cyclical Reflection Token Scheduling

ICLR 2026poster

Large reasoning models (LRMs), such as OpenAI's o1 and DeepSeek-R1, harness test-time scaling to perform multi-step reasoning for complex problem-solving. This reasoning process, executed before producing final answers, is often guided by special juncture tokens that prompt self-evaluative reflectio…

Cited by 0SourcecodeScholar