← Search

Navid Akhavan Attar

2 accepted papers

2026

DirMoE: Dirichlet-Routed Mixture of Experts

ICLR 2026poster

Mixture-of-Experts (MoE) models have demonstrated exceptional performance in large-scale language models. Existing routers typically rely on non-differentiable Top-$k$+Softmax, limiting their performance and scalability. We argue that two distinct decisions, which experts to activate and how to dist…

Cited by 0SourceScholar
2026

Softmax is not Enough (for Adaptive Conformal Classification)

ICLR 2026poster

The merit of Conformal Prediction (CP), as a distribution-free framework for uncertainty quantification, depends on generating prediction sets that are efficient, reflected in small average set sizes, while adaptive, meaning they signal uncertainty by varying in size according to input difficulty. A…

Cited by 0SourceScholar