← Search

Timon Klein

2 accepted papers

2026

Tucker Attention: A generalization of approximate attention mechanisms

ICML 2026poster

The pursuit of reducing the memory footprint of the self-attention mechanism in multi-headed self attention (MHA) spawned a rich portfolio of methods, e.g., group-query attention (GQA) and multi-head latent attention (MLA). The methods leverage specialized low-rank factorizations across embedding di…

Cited by 0SourceScholar
2025

A geometric framework for momentum-based optimizers for low-rank training

NeurIPS 2025poster

Low-rank pre-training and fine-tuning have recently emerged as promising techniques for reducing the computational and storage costs of large neural networks. Training low-rank parameterizations typically relies on conventional optimizers such as heavy ball momentum methods or Adam. In this work, we…

Cited by 0SourceScholar