2026
Pretraining with hierarchical memories: separating long-tail and common knowledge
ICLR 2026poster
The impressive performance gains of modern language models currently rely on scaling parameters: larger models store more world knowledge and reason better. Yet compressing all world knowledge into parameters is unnecessary, as only a fraction is used per prompt, and impractical for edge devices wit…