2026
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
ICLR 2026poster
Large Language Models (LLMs) achieve strong performance on diverse tasks but often exhibit cognitive inertia, struggling to follow instructions that conflict with the standardized patterns learned during supervised fine-tuning (SFT). To evaluate this limitation, we propose Inverse IFEval, a benchmar…