2025
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
NeurIPS 2025poster
The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contamination part, or prompt, functioning as a new, trainable expert. Despite its populari…