2026
Inner-layer Token Self-Modulation as Another Scaling Axis for LLMs
ICML 2026poster
LLMs have traditionally scaled along dense dimensions, where performance is coupled with near-linear increases in computational cost. While MoE decouples capacity from compute, it introduces large memory overhead and hardware efficiency challenges. To overcome these, we propose token-indexed paramet…