← Search

Malik Khalaf

1 accepted papers

2026

QKV Projections Require a Fraction of Their Memory

ICLR 2026poster

The Multi-Head Attention mechanism is central to LLM operation, and multiple works target its compute and memory efficiency during training. While most works focus on approximating the scaled dot product, the memory consumption of the linear projections that compute the $Q$, $K$, and $V$ tensors fro…

Cited by 0SourceScholar