← Search

Thiziri Nait Saada

3 accepted papers

2026

Removing Noise, not Finding Gold: Quality Filtering for Large-Scale Pretraining

ICML 2026poster

Large-scale models are pretrained on massive web-crawled datasets containing documents of mixed quality, making data filtering essential. A popular method is Classifier-based Quality Filtering (CQF), which trains a binary classifier to distinguish between pretraining data and a small, high-quality s…

Cited by 0SourceScholar
2025

Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention Layers

ICML 2025poster

Attention layers are the core component of transformers, the current state-of-the-art neural network architecture. Alternatives to softmax-based attention are being explored due to its tendency to hinder effective information flow. Even *at initialisation*, it remains poorly understood why the propa…

Cited by 0SourcePDFScholar
2024

Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian Processes

ICLR 2024poster

The infinitely wide neural network has been proven a useful and manageable mathematical model that enables the understanding of many phenomena appearing in deep learning. One example is the convergence of random deep networks to Gaussian processes that enables a rigorous analysis of the way the choi…

Cited by 1SourcePDFScholar