2026
TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
ICML 2026poster
Mixture-of-Experts (MoE) models achieve remarkable performance by sparsely activating specialized experts, yet their massive parameters in experts pose significant challenges for deployment. While low-rank quantization offers a promising route to compress MoE models, existing methods still incur non…