2026
Bring Future Vision: Dynamic Computation Allocation Guided by Lightweight Feature Forecaster
ICML 2026poster
The deployment of large language models (LLMs) in real-world applications is increasingly limited by their high inference cost. While recent advances in dynamic token-level computation allocation attempt to improve efficiency by selectively activating model components per token, existing methods rel…