← Search

Mark Rofin

3 accepted papers

2026

Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors

ICLR 2026poster

Trained Transformers have been shown to compute abstract features that appear redundant for predicting the immediate next token. We identify which components of the gradient signal from the next-token prediction objective give rise to this phenomenon, and we propose a method to estimate the influenc…

Cited by 0SourcecodeScholar
2025

Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers

ICML 2025poster

Chain-of-thought reasoning and scratchpads have emerged as critical tools for enhancing the computational capabilities of transformers. While theoretical results show that polynomial-length scratchpads can extend transformers' expressivity from $TC^0$ to $PTIME$, their required length remains poorly…

Cited by 2SourcePDFScholar