← Search

Mike Lasby

3 accepted papers

2026

REAP the Experts: Why Pruning Prevails for One-Shot MoE compression

ICLR 2026poster

Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating research into expert compression. Contrary to recent findings favouring expert *merging* on discriminative benchmarks, we f…

Cited by 0SourcecodeScholar
2024

Dynamic Sparse Training with Structured Sparsity

ICLR 2024poster

Dynamic Sparse Training (DST) methods achieve state-of-the-art results in sparse neural network training, matching the generalization of dense models while enabling sparse training and inference. Although the resulting models are highly sparse and theoretically less computationally expensive, achiev…

2024

Navigating Extremes: Dynamic Sparsity in Large Output Spaces

NeurIPS 2024poster

In recent years, Dynamic Sparse Training (DST) has emerged as an alternative to post-training pruning for generating efficient models. In principle, DST allows for a much more memory efficient training process, as it maintains sparsity throughout the entire training run. However, current DST implem…

Cited by 2SourcePDFScholar