2026
UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning
ICLR 2026poster
While Mixture of Experts (MoE) models achieve remarkable efficiency by activating only subsets of parameters, they suffer from high memory access costs during inference. Memory-layer architectures offer an appealing alternative with very few memory access, but previous attempts like UltraMem have on…