← Search

Maty Bohacek

2 accepted papers

2026

Uncovering Competency Gaps in Large Language Models and Their Benchmarks

ICML 2026poster

The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can obscure (i) particular sub-areas where the models are weak ("model gaps") (ii) imbalanced coverage in the benchmarks themselves ("benchmark gaps"). To automa…

Cited by 0SourceScholar
2026

Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders

ICLR 2026poster

Despite their impressive performance, generative image models trained on large-scale datasets frequently fail to produce images with seemingly simple concepts -- e.g., human hands or objects appearing in groups of four -- that are reasonably expected to appear in the training data. These failure mod…

Cited by 0SourceScholar