← Search

David Grangier

27 accepted papers

2026

Optimal Splitting of Language Models from Mixtures to Specialized Domains

ICML 2026poster

Language models achieve impressive performance on a variety of knowledge, language, and reasoning tasks due to the scale and diversity of pretraining data available. The standard training recipe is a two-stage paradigm: pretraining first on the full corpus of data followed by specialization on a muc…

Cited by 0SourceScholar
2026

Pretraining with hierarchical memories: separating long-tail and common knowledge

ICLR 2026poster

The impressive performance gains of modern language models currently rely on scaling parameters: larger models store more world knowledge and reason better. Yet compressing all world knowledge into parameters is unnecessary, as only a fraction is used per prompt, and impractical for edge devices wit…

Cited by 0SourceScholar
2026

Removing Noise, not Finding Gold: Quality Filtering for Large-Scale Pretraining

ICML 2026poster

Large-scale models are pretrained on massive web-crawled datasets containing documents of mixed quality, making data filtering essential. A popular method is Classifier-based Quality Filtering (CQF), which trains a binary classifier to distinguish between pretraining data and a small, high-quality s…

Cited by 0SourceScholar
2025

Assessing the Role of Data Quality in Training Bilingual Language Models

EMNLP 2025

Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly between languages as prior works show that adding more languages can degrade performance for some languages (such as English

2025

No Need to Talk: Asynchronous Mixture of Language Models

ICLR 2025spotlight

We introduce SMALLTALK LM, an innovative method for training a mixture of language models in an almost asynchronous manner. Each model of the mixture specializes in distinct parts of the data distribution, without the need of high-bandwidth communication between the nodes training each model. At inf…

Cited by 1SourcePDFScholar
2025

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

ICML 2025poster

A widespread strategy to obtain a language model that performs well on a target domain is to finetune a pretrained model to perform unsupervised next-token prediction on data from that target domain. Finetuning presents two challenges: \textit{(i)} if the amount of target data is limited, as in most…

Cited by 1SourcePDFScholar
2025

Scaling Laws for Optimal Data Mixtures

NeurIPS 2025poster

Large foundation models are typically trained on data from multiple domains, with the data mixture--the proportion of each domain used--playing a critical role in model performance. The standard approach to selecting this mixture relies on trial and error, which becomes impractical for large-scale p…

Cited by 0SourceScholar
2025

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging

ICML 2025poster

Machine learning models are routinely trained on a mixture of different data domains. Different domain weights yield very different downstream performances. We propose the Soup-of-Experts, a novel architecture that can instantiate a model at test time for any domain weights with minimal computation…

Cited by 0SourcePDFScholar
2025

Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

ICLR 2025poster

Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain these models is only available in limited amount for most tasks. In this work, we build specialist models from large gen…

Cited by 3SourcePDFScholar
2025

Training Bilingual LMs with Data Constraints in the Targeted Language

ACL 2025finding

Large language models are trained on massive scrapes of the web, as required by current scaling laws. Most progress is made for English, given its abundance of high-quality pretraining data. For most other languages, however, such high quality pretraining data is unavailable. In this work, we study…

2024

Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP

NeurIPS 2024poster

Large pretrained vision-language models like CLIP have shown promising generalization capability, but may struggle in specialized domains (e.g., satellite imagery) or fine-grained classification (e.g., car models) where the visual concepts are unseen or under-represented during pretraining. Prompt l…

Cited by 0SourcePDFScholar
2024

Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling

ACL 2024long

Large language models are trained on massive scrapes of the web, which are often unstructured, noisy, and poorly phrased. Current scaling laws show that learning from such data requires an abundance of both compute and data, which grows with the size of the model being trained. This is infeasible bo…

Cited by 57SourcePDFScholar
2022

A Natural Diet: Towards Improving Naturalness of Machine Translation Output

ACL 2022findings

Machine translation (MT) evaluation often focuses on accuracy and fluency, without paying much attention to translation style. This means that, even when considered accurate and fluent, MT output can still sound less natural than high quality human translations or text originally written in the targ…

Cited by 18SourcePDFScholar
2022

Learning Strides in Convolutional Neural Networks

ICLR 2022oral

Convolutional neural networks typically contain several downsampling operators, such as strided convolutions or pooling layers, that progressively reduce the resolution of intermediate representations. This provides some shift-invariance while reducing the computational complexity of the whole archi…

2022

On Systematic Style Differences between Unsupervised and Supervised MT and an Application for High-Resource Machine Translation

NAACL 2022long

Modern unsupervised machine translation (MT) systems reach reasonable translation quality under clean and controlled data conditions. As the performance gap between supervised and unsupervised MT narrows, it is interesting to ask whether the different training methods result in systematically differ…

2021

AUXILIARY TASK UPDATE DECOMPOSITION: THE GOOD, THE BAD AND THE NEUTRAL

ICLR 2021poster

While deep learning has been very beneficial in data-rich settings, tasks with smaller training set often resort to pre-training or multitask learning to leverage data from other tasks. In this case, careful consideration is needed to select tasks and model parameterizations such that updates from t…

Cited by 30SourcecodeScholar
2021

Learning From Heterogeneous Eeg Signals with Differentiable Channel Reordering

ICASSP 2021accepted

We propose CHARM, a method for training a single neural network across inconsistent input channels. Our work is motivated by Electroencephalography (EEG), where data collection protocols from different headsets result in varying channel ordering and number, which limits the feasibility of transferri…

Cited by 0SourceScholar
2019

3D Human Pose Estimation in Video With Temporal Convolutions and Semi-Supervised Training

CVPR 2019poster

In this work, we demonstrate that 3D poses in video can be effectively estimated with a fully convolutional model based on dilated temporal convolutions over 2D keypoints. We also introduce back-projection, a simple and effective semi-supervised training method that leverages unlabeled video data. W…

Cited by 1464PDFcodeScholar
2018

Analyzing Uncertainty in Neural Machine Translation

ICML 2018oral

Machine translation is a popular test bed for research in neural sequence-to-sequence models but despite much recent research, there is still a lack of understanding of these models. Practitioners report performance degradation with large beams, the under-estimation of rare words and a lack of diver…

2017

Convolutional Sequence to Sequence Learning

ICML 2017poster

The prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks. Compared to recurrent models, computations over all elements can be fully p…

2017

Efficient Softmax Approximation for GPUs

ICLR 2017workshop

We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size by exploiting the unbalanced word distribution to form clusters that explicitly m…

Cited by 348SourceScholar
2017

Efficient softmax approximation for GPUs

ICML 2017poster

We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size by exploiting the unbalanced word distribution to form clusters that explicitly m…