ICLR 2026poster0 citations

RMAAT: Astrocyte-Inspired Memory Compression and Replay for Efficient Long-Context Transformers

Md Zesun Ahmed Mia, Malyaban Bal, Abhronil Sengupta

Abstract

The quadratic complexity of self-attention mechanism presents a significant impediment to applying Transformer models to long sequences. This work explores computational principles derived from astrocytes—glial cells critical for biological memory and synaptic modulation—as a complementary approach to conventional architectural modifications for efficient self-attention. We introduce the Recurrent Memory Augmented Astromorphic Transformer (RMAAT), an architecture integrating abstracted astrocyte functionalities. RMAAT employs a recurrent, segment-based processing strategy where persistent memory tokens propagate contextual information. An adaptive compression mechanism, governed by a novel retention factor derived from simulated astrocyte long-term plasticity (LTP), modulates these tokens. Attention within segments utilizes an efficient, linear-complexity mechanism inspired by astrocyte short-term plasticity (STP). Training is performed using Astrocytic Memory Replay Backpropagation (AMRB), a novel algorithm designed for memory efficiency in recurrent networks. Evaluations on the Long Range Arena (LRA) benchmark demonstrate RMAAT's competitive accuracy and substantial improvements in computational and memory efficiency, indicating the potential of incorporating astrocyte-inspired dynamics into scalable sequence models.

Brain-inspired machine learningAstromorphic transformersShort-term PlasticityLong-term PLasticityLong-context sequence modeling
BibTeX
@inproceedings{
mia2026rmaat,
title={{RMAAT}: Astrocyte-Inspired Memory Compression and Replay for Efficient Long-Context~Transformers},
author={Md Zesun Ahmed Mia and Malyaban Bal and Abhronil Sengupta},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=sTkJdbVxsI}
}
RMAAT: Astrocyte-Inspired Memory Compression and Replay for Efficient Long-Context Transformers · ICLR 2026