2026
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
AAAI 2026technical
Mixture-of-Experts (MoE) architectures decouple model capacity from per-token computation, enabling scaling beyond the computational limits imposed by dense scaling laws. Yet how MoE architectures shape knowledge acquisition during pre-training—and how this process differs from dense architectures—r