2026
Mapping Overlaps in Benchmarks through Perplexity in the Wild
ICLR 2026poster
We construct benchmark signatures that capture the capacity required for strong performance to characterize large language model (LLM) benchmarks and their meaningful overlaps. Formally, we define them as sets of salient tokens drawn from **in-the-wild** corpora whose LLM token perplexity, reflectin…