2026
LatentLLM: Activation-Aware Transform to Multi-Head Latent Attention
AAAI 2026technical
Modern foundation models such as large language models (LLMs) require a massive amount of computational and memory resources. We propose a new framework to convert such LLMs into a reduced-dimension latent structure. Our method extends a local activation-aware tensor decomposition to a global attent