2025
Training Medical QA Models Based on Mixed Rewards from Multiple-Choice and Open-Ended Questions
EMNLP 2025
Reinforcement learning (RL) for large language models (LLMs) typically requires clear reward signals, which are often unavailable for open-ended (OE) questions where answer evaluation is ambiguous without scalable expert labeling. We investigate whether LLMs benefit from training on mixed data with