← Search

Elvis Dohmatob

25 accepted papers

2026

Why Less is More (Sometimes): A Theory of Data Curation

ICLR 2026poster

This paper introduces a theoretical framework to resolve a central paradox in modern machine learning: When is it better to use less data? This question has become critical as classical scaling laws suggesting ``more is more'' (Sun et al., 2025) are challenged by methods like LIMO (``less is more'')…

Cited by 0SourcecodeScholar
2025

Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification

ICLR 2025poster

Large Language Models (LLM) are increasingly trained on data generated by other LLMs, either because generated text and images become part of the pre-training corpus, or because synthetized data is used as a replacement for expensive human-annotation. This raises concerns about *model collapse*, a d…

Cited by 3SourcePDFScholar
2025

Improving the Scaling Laws of Synthetic Data with Deliberate Practice

ICML 2025oral

Inspired by the principle of deliberate practice in human learning, we propose Deliberate Practice for Synthetic Data Generation (DP), a novel framework that improves sample efficiency through dynamic synthetic data generation. Prior work has shown that scaling synthetic data is inherently challengi…

Cited by 0SourcePDFScholar
2025

The Pitfalls of Memorization: When Memorization Hurts Generalization

ICLR 2025poster

Neural networks often learn simple explanations that fit the majority of the data while memorizing exceptions that deviate from these explanations. This behavior leads to poor generalization when the learned explanations rely on spurious correlations. In this work, we formalize $\textit{the interpla…

2024

A Tale of Tails: Model Collapse as a Change of Scaling Laws

ICML 2024poster

As AI model size grows, neural *scaling laws* have become a crucial tool to predict the improvements of large models when increasing capacity and the size of original (human or natural) training data. Yet, the widespread use of popular models means that the ecosystem of online data and text will co-…

Cited by 58SourcePDFScholar
2023

Contextual bandits with concave rewards, and an application to fair ranking

ICLR 2023poster

We consider Contextual Bandits with Concave Rewards (CBCR), a multi-objective bandit problem where the desired trade-off between the rewards is defined by a known concave objective function, and the reward vector depends on an observed stochastic context. We present the first algorithm with provably…

Cited by 4SourcePDFScholar
2022

Scalable MCMC Sampling for Nonsymmetric Determinantal Point Processes

ICML 2022oral

A determinantal point process (DPP) is an elegant model that assigns a probability to every subset of a collection of $n$ items. While conventionally a DPP is parameterized by a symmetric kernel matrix, removing this symmetry constraint, resulting in nonsymmetric DPPs (NDPPs), leads to significant i…

2022

Scalable Sampling for Nonsymmetric Determinantal Point Processes

ICLR 2022spotlight

A determinantal point process (DPP) on a collection of $M$ items is a model, parameterized by a symmetric kernel matrix, that assigns a probability to every subset of those items. Recent work shows that removing the kernel symmetry constraint, yielding nonsymmetric DPPs (NDPPs), can lead to signifi…

2021

Scalable Learning and MAP Inference for Nonsymmetric Determinantal Point Processes

ICLR 2021oral

Determinantal point processes (DPPs) have attracted significant attention in machine learning for their ability to model subsets drawn from a large item collection. Recent work shows that nonsymmetric DPP (NDPP) kernels have significant advantages over symmetric kernels in terms of modeling power an…

2020

Learning disconnected manifolds: a no GAN’s land

ICML 2020poster

Typical architectures of Generative Adversarial Networks make use of a unimodal latent/input distribution transformed by a continuous generator. Consequently, the modeled distribution always has connected support which is cumbersome when learning a disconnected set of manifolds. We formalize this pr…

Cited by 48SourcePDFScholar
2020

On the Convergence of Smooth Regularized Approximate Value Iteration Schemes

NeurIPS 2020spotlight

Entropy regularization, smoothing of Q-values and neural network function approximator are key components of the state-of-the-art reinforcement learning (RL) algorithms, such as Soft Actor-Critic~\cite{haarnoja2018soft}. Despite the widespread use, the impact of these core techniques on the converge…

Cited by 10SourcePDFScholar
2019

Learning Nonsymmetric Determinantal Point Processes

NeurIPS 2019poster

Determinantal point processes (DPPs) have attracted substantial attention as an elegant probabilistic model that captures the balance between quality and diversity within sets. DPPs are conventionally parameterized by a positive semi-definite kernel matrix, and this symmetric kernel encodes only re…

2016

Learning brain regions via large-scale online structured sparse dictionary learning

NeurIPS 2016poster

We propose a multivariate online dictionary-learning method for obtaining decompositions of brain images with structured and sparse components (aka atoms). Sparsity is to be understood in the usual sense: the dictionary atoms are constrained to contain mostly zeros. This is imposed via an $\ell_1$-n…

Cited by 23SourcePDFScholar
2016

Local Q-linear convergence and finite-time active set identification of ADMM on a class of penalized regression problems

ICASSP 2016accepted

We study the convergence of the ADMM (Alternating Direction Method of Multipliers) algorithm on a broad range of penalized regression problems including the Lasso, Group-Lasso and Graph-Lasso,(isotropic) TV-L1, Sparse Variation, and others. First, we establish a fixed-point iterationvia a nonlinear…

Cited by 0SourceScholar