2026
The Consistency Trap in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
ICML 2026poster
Large language models are often evaluated for correctness on isolated questions. But modern deployments also rely on a different property: whether the model stays consistent as it generates, critiques, and revises over multiple steps that rely on the same underlying concepts. In these settings, *sel…