2026
FrugalRAG: Less is More in RL Finetuning for Multi-hop Question Answering
ICLR 2026poster
Reinforcement learning (RL) based on the final answer's reward has driven recent progress in small language models (SLMs) on reasoning-heavy tasks such as math and code. However, applying the same techniques to retrieval-augmented generation (RAG) benchmarks like multi-hop QA has yielded limited gai…