2026
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
ICLR 2026poster
Evaluation using Large Language Model (LLM) judges has been widely adopted in English and shown to be effective for automatic evaluation. However, their performance does not generalize well to non-English settings, and it remains unclear what constitutes effective multilingual training for such judg…