2026
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
ICML 2026poster
Large Language Models (LLMs) demonstrate remarkable capabilities but face deployment challenges due to their high computational demands. Traditional pruning methods reduce these costs by permanently removing parameters, which inevitably leads to performance degradation. To mitigate this issue, we pr…