2026
Understanding Dynamic Compute Allocation in Recurrent Transformers
ICML 2026poster
Token-level adaptive computation seeks to reduce inference cost by allocating more computation to harder tokens and less to easier ones. However, prior work is primarily evaluated on natural-language benchmarks using task-level metrics, where token-level difficulty is unobservable and confounded wit…