← Search

Alexa Tartaglini

1 accepted papers

2026

Outcome-Based Rewards Do Not Guarantee Faithful and Verifiable Reasoning

ICML 2026poster

Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that the reasoning chains trained through RLVR represent how a model gets to its answer. In this paper, we develop two metric…

Cited by 0SourceScholar