← Search

Callum Stuart McDougall

3 accepted papers

2026

What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation

ICLR 2026poster

Prior work suggests that language models, while trained on next token prediction, show implicit planning behavior: they may select the next token in preparation to a predicted future token, such as a likely rhyming word, as supported by a prior qualitative study of Claude 3.5 Haiku using a cross-lay…

Cited by 0SourceScholar
2025

Internal states before wait modulate reasoning patterns

EMNLP 2025

Prior work has shown that a significant driver of performance in reasoning models is their ability to reason and self-correct. A distinctive marker in these reasoning traces is the token wait , which often signals reasoning behavior such as backtracking. Despite being such a complex behavior, little

Cited by 0SourcePDFScholar
2025

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

ICML 2025poster

Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most prior work evaluates progress using unsupervised proxy metrics with unclear practical relevance. We introduce SAEBench, a c…

Cited by 0SourcePDFScholar