← Search

Suhaas Bhat

2 accepted papers

2026

Building Reliable Long-Form Generation via Hallucination Rejection Sampling

ICML 2026poster

Large language models (LLMs) have achieved remarkable progress in open-ended text generation, yet they remain prone to hallucinating incorrect or unsupported content, which undermines their reliability. This issue is exacerbated in long-form generation due to hallucination snowballing, a phenomenon …

Cited by 0SourceScholar
2026

Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Creative Writing

ICML 2026poster

Reinforcement learning with verifiable rewards (RLVR) on foundation models has led to significant improvements in math and code generation. Extending these gains to open-ended domains remains challenging: ground-truth verification is unavailable, human annotation is expensive, and learnt reward mode…

Cited by 0SourceScholar