← Search

Amirhossein Vahidi

2 accepted papers

2026

DirMoE: Dirichlet-Routed Mixture of Experts

ICLR 2026poster

Mixture-of-Experts (MoE) models have demonstrated exceptional performance in large-scale language models. Existing routers typically rely on non-differentiable Top-$k$+Softmax, limiting their performance and scalability. We argue that two distinct decisions, which experts to activate and how to dist…

Cited by 0SourceScholar
2024

Probabilistic Self-supervised Representation Learning via Scoring Rules Minimization

ICLR 2024poster

% Self-supervised learning methods have shown promising results across a wide range of tasks in computer vision, natural language processing, and multimodal analysis. However, self-supervised approaches come with a notable limitation, dimensional collapse, where a model doesn't fully utilize its cap…