← Search

Pedram Akbarian

7 accepted papers

2025

On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts

NeurIPS 2025poster

The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contamination part, or prompt, functioning as a new, trainable expert. Despite its populari…

Cited by 0SourceScholar
2025

Statistical Advantages of Perturbing Cosine Router in Mixture of Experts

ICLR 2025poster

The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable performance in image and language tasks and exhibits better ability to mitigate the representation collapse issue, which…

Cited by 6SourcePDFScholar
2025

Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts

AISTATS 2025poster

We conduct the convergence analysis of parameter estimation in the contaminated mixture of experts. This model is motivated from the prompt learning problem where ones utilize prompts, which can be formulated as experts, to fine-tune a large-scale pre-trained model for learning downstream tasks. The…

Cited by 0SourceScholar
2024

A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts

ICML 2024poster

Mixture-of-experts (MoE) model incorporates the power of multiple submodels via gating functions to achieve greater performance in numerous regression and classification applications. From a theoretical perspective, while there have been previous attempts to comprehend the behavior of that model und…

Cited by 8SourcePDFScholar
2024

Improving Computational Complexity in Statistical Models with Local Curvature Information

ICML 2024poster

It is known that when the statistical models are singular, i.e., the Fisher information matrix at the true parameter is degenerate, the fixed step-size gradient descent algorithm takes polynomial number of steps in terms of the sample size $n$ to converge to a final statistical radius around the tru…

Cited by 0SourcePDFScholar
2024

Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts

ICLR 2024poster

Top-K sparse softmax gating mixture of experts has been widely used for scaling up massive deep-learning architectures without increasing the computational cost. Despite its popularity in real-world applications, the theoretical understanding of that gating function has remained an open problem. The…

Cited by 17SourcePDFScholar