Representational Curvature Shapes Behavioral Uncertainty in Large Language Models
In autoregressive large language models (LLMs), temporal straightening offers an account of how the next-token prediction objective shapes representations. Across layers, models progressively straighten the trajectory of input sequences in activation space, potentially facilitating extrapolation to …