← Search

Matvey Kairov

1 accepted papers

2026

GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent

ICML 2026poster

Many large language model applications require conditioning on long contexts. Transformers typically support this by storing a large per-layer KV-cache of past activations, which incurs substantial memory overhead. A desirable alternative is compressive memory: read a context once, store it in a com…

Cited by 0SourceScholar