← Search

Itamar Rocha Filho

1 accepted papers

2026

Parameter-Efficient Reinforcement Learning using Prefix Optimization

ICLR 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) is a leading approach for tuning language models on mathematical reasoning tasks. However, it remains unclear whether RLVR's gains stem from genuine reasoning improvements or simply from steering the model toward answer formats that already appea…

Cited by 0SourcecodeScholar