2026
A$^2$Search: Ambiguity-Aware Question Answering with Reinforcement Learning
ICLR 2026poster
Recent advances in Large Language Models (LLMs) and Reinforcement Learning (RL) have led to strong performance in open-domain question answering (QA). However, existing models still struggle with questions that admit multiple valid answers. Standard QA benchmarks, which typically assume a single gol…