AAAI 2026technical0 citations

LatentLLM: Activation-Aware Transform to Multi-Head Latent Attention

Toshiaki Koike-Akino, Xiangyu Chen, Jing Liu, Ye Wang, Pu (Perry) Wang, Matthew Brand

Abstract

Modern foundation models such as large language models (LLMs) require a massive amount of computational and memory resources. We propose a new framework to convert such LLMs into a reduced-dimension latent structure. Our method extends a local activation-aware tensor decomposition to a global attention-aware joint tensor decomposition. Our framework can significantly improve the model accuracy over the existing model compression methods when reducing the latent dimension to realize computationally/memory-efficient LLMs. We show the benefit on several benchmark including multi-modal reasoning tasks.

BibTeX
@inproceedings{aaai2026_latentllmactivat,
  title = {LatentLLM: Activation-Aware Transform to Multi-Head Latent Attention},
  author = {Toshiaki Koike-Akino and Xiangyu Chen and Jing Liu and Ye Wang and Pu (Perry) Wang and Matthew Brand},
  booktitle = {AAAI 2026},
  year = {2026}
}
LatentLLM: Activation-Aware Transform to Multi-Head Latent Attention · AAAI 2026