← Search

Nurit Cohen Inger

1 accepted papers

2025

Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon

EMNLP 2025

Large language models (LLMs) often appear to excel on public benchmarks, but these high scores may mask an overreliance on dataset-specific surface cues rather than true language understanding. We introduce the **Chameleon Benchmark Overfit Detector (C-BOD)**, a meta-evaluation framework designed to