← Search

Alex Bie

7 accepted papers

2025

Escaping Collapse: The Strength of Weak Data for Large Language Model Training

NeurIPS 2025poster

Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown that without proper curation it can cause LLM performance to plateau, or even "collapse", after many training iteration…

Cited by 0SourceScholar
2024

Private prediction for large-scale synthetic text generation

EMNLP 2024finding

We present an approach for generating differentially private synthetic text using large language models (LLMs), via private prediction. In the private prediction framework, we only require the output synthetic data to satisfy differential privacy guarantees. This is in contrast to approaches that tr…

2023

Private Distribution Learning with Public Data: The View from Sample Compression

NeurIPS 2023spotlight

We study the problem of private distribution learning with access to public data. In this setup, which we refer to as *public-private learning*, the learner is given public and private samples drawn from an unknown distribution $p$ belonging to a class $\mathcal Q$, with the goal of outputting an es…

Cited by 21SourcePDFScholar
2021

Don’t Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence

NeurIPS 2021poster

Although machine learning models trained on massive data have led to breakthroughs in several areas, their deployment in privacy-sensitive domains remains limited due to restricted access to data. Generative models trained with privacy constraints on private data can sidestep this challenge, providi…

Cited by 83SourcePDFScholar