← Search

Zhangyang Peng

1 accepted papers

2026

OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale

ICML 2026poster

Mixture-of-Experts (MoE) architectures are evolving towards finer granularity to improve parameter efficiency. However, existing MoE designs face an inherent trade-off between the granularity of expert specialization and hardware execution efficiency. In this paper, we propose OmniMoE, a system-algo…

Cited by 0SourceScholar