← Search

C Thomas

1 accepted papers

2026

Pretraining with hierarchical memories: separating long-tail and common knowledge

ICLR 2026poster

The impressive performance gains of modern language models currently rely on scaling parameters: larger models store more world knowledge and reason better. Yet compressing all world knowledge into parameters is unnecessary, as only a fraction is used per prompt, and impractical for edge devices wit…

Cited by 0SourceScholar