← Search

Daria Soboleva

3 accepted papers

2025

Power Lines: Scaling laws for weight decay and batch size in LLM pre-training

NeurIPS 2025poster

Efficient LLM pre-training requires well-tuned hyperparameters (HPs), including learning rate η and weight decay λ. We study scaling laws for HPs: formulas for how to scale HPs as we scale model size N, dataset size D, and batch size B. Recent work suggests the AdamW timescale, τ = B/(ηλD), should…

Cited by 0SourceScholar
2025

Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs

ICLR 2025poster

LLMs are commonly trained with a learning rate (LR) warmup, followed by cosine decay to 10% of the maximum (10x decay). In a large-scale empirical study, we show that under an optimal peak LR, a simple linear decay-to-zero (D2Z) schedule consistently outperforms other schedules when training at comp…

Cited by 2SourcePDFScholar
2021

Replacing Human Audio with Synthetic Audio for on-Device Unspoken Punctuation Prediction

ICASSP 2021accepted

We present a novel multi-modal unspoken punctuation prediction system for the English language which combines acoustic and text features. We demonstrate for the first time, that by relying exclusively on synthetic data generated using a prosody-aware text-to-speech system, we can outperform a model…

Cited by 0SourceScholar