2025
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
EMNLP 2025
Large language models (LLMs) demonstrate proficiency across numerous computational tasks, yet their inner workings remain unclear. In theory, the combination of causal self-attention and multilayer perceptron allows every token to access and compute information based on all preceding tokens. In prac