← Search

Benjamin Thérien

4 accepted papers

2026

$\mu$LO: Compute-Efficient Meta-Generalization of Learned Optimizers

ICLR 2026poster

Learned optimizers (LOs) have the potential to significantly reduce the wall-clock training time of neural networks. However, they can struggle to optimize unseen tasks (*meta-generalize*), especially when training networks wider than those seen during meta-training. To address this, we derive the M…

Cited by 0SourcecodeScholar
2026

MuLoCo: Muon is a Practical Inner Optimizer for DiLoCo

ICML 2026poster

DiLoCo is a powerful framework for training large language models (LLMs) under networking constraints, allowing for increased parallelism and accelerator utilization in data center settings. A critical but often overlooked factor in DiLoCo’s behavior is the choice of inner optimizer, which shapes th…

Cited by 0SourceScholar
2025

Dense Backpropagation Improves Training for Sparse Mixture-of-Experts

NeurIPS 2025poster

Mixture of Experts (MoE) pretraining is more scalable than dense Transformer pretraining, because MoEs learn to route inputs to a sparse set of their feedforward parameters. However, this means that MoEs only receive a sparse backward update, leading to training instability and suboptimal performanc…

Cited by 0SourcecodeScholar
2022

Parametric Scattering Networks

CVPR 2022oral

The wavelet scattering transform creates geometric invariants and deformation stability. In multiple signal domains, it has been shown to yield more discriminative representations compared to other non-learned representations and to outperform learned representations in certain tasks, particularly o…

Cited by 26PDFcodeScholar