2026
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
ICLR 2026poster
Verifiers—functions assigning rewards to agent behavior—have been key for AI progress in domains such as math, code and games. However, extending these gains to domains without clear-cut success criteria (e.g., computer use) remains a challenge: while humans can recognize suitable outcomes, translat…