← Search

Rizhen Hu

3 accepted papers

2026

Grouter: Decoupling Routing from Representation for Accelerated MoE Training

ICML 2026poster

Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneously train expert weights while searching for an optimal routing policy within a vast combinatorial space. This entanglement often leads to sluggish convergenc…

Cited by 0SourceScholar
2026

Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization

ICML 2026poster

Sparse Mixture-of-Experts (MoE) models scale Transformers efficiently but suffer from expert overlap, where different experts process similar tokens and learn redundant functions, resulting in ambiguous routing and underutilized capacity. While architectural solutions like DeepSeek-style shared expe…

Cited by 0SourceScholar
2025

MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization

NeurIPS 2025poster

As distributed optimization scales to meet the demands of Large Language Model (LLM) training, hardware failures become increasingly non-negligible. Existing fault-tolerant training methods often introduce significant computational or memory overhead, demanding additional resources. To address this…

Cited by 0SourceScholar