2026
Unified Episodic and Semantic Memory via Modulating Transformer FeedForward Layers
ICML 2026poster
It is widely recognized that, after generative pre-training, Transformer FeedForward layers implicitly function as semantic memory, encoding linguistic and factual knowledge, while the contexts in key–value (KV) cache contain raw events, serving as the source of models' episodic memory. In this work…