← Search

Giannis Daras

21 accepted papers

2026

Ambient Dataloops: Generative Models for Dataset Refinement

ICML 2026poster

We propose Ambient Dataloops, an iterative framework for refining datasets that makes it easier for diffusion models to learn the underlying data distribution. Modern datasets contain samples of highly varying quality, and training directly on such heterogeneous data often yields suboptimal models. …

Cited by 0SourceScholar
2025

Ambient Diffusion Omni: Training Good Models with Bad Data

NeurIPS 2025spotlight

We show how to use low-quality, synthetic, and out-of-distribution images to improve the quality of a diffusion model. Typically, diffusion models are trained on curated datasets that emerge from highly filtered data pools from the Web and other sources. We show that there is immense value in the lo…

Cited by 0SourcecodeScholar
2025

Ambient Diffusion Posterior Sampling: Solving Inverse Problems with Diffusion Models Trained on Corrupted Data

ICLR 2025poster

We provide a framework for solving inverse problems with diffusion models learned from linearly corrupted data. Firstly, we extend the Ambient Diffusion framework to enable training directly from measurements corrupted in the Fourier domain. Subsequently, we train diffusion models for MRI with acces…

2025

Ambient Proteins - Training Diffusion Models on Noisy Structures

NeurIPS 2025spotlight

We present Ambient Protein Diffusion, a framework for training protein diffusion models that generates structures with unprecedented diversity and quality. State-of-the-art generative models are trained on computationally derived structures from AlphaFold2 (AF), as experimentally determined structur…

Cited by 0SourceScholar
2025

Does Generation Require Memorization? Creative Diffusion Models using Ambient Diffusion

ICML 2025poster

There is strong empirical evidence that the stateof-the-art diffusion modeling paradigm leads to models that memorize the training set, especially when the training set is small. Prior methods to mitigate the memorization problem often lead to decrease in image quality. Is it possible to obtain stro…

Cited by 0SourcePDFScholar
2025

How Much is a Noisy Image Worth? Data Scaling Laws for Ambient Diffusion.

ICLR 2025poster

The quality of generative models depends on the quality of the data they are trained on. Creating large-scale, high-quality datasets is often expensive and sometimes impossible, e.g.~in certain scientific applications where there is no access to clean data due to physical or instrumentation constrai…

2025

Infilling Score: A Pretraining Data Detection Algorithm for Large Language Models

ICLR 2025poster

In pretraining data detection, the goal is to detect whether a given sentence is in the dataset used for training a Large Language Model LLM). Recent methods (such as Min-K % and Min-K%++) reveal that most training corpora are likely contaminated with both sensitive content and evaluation benchmarks…

Cited by 0SourcePDFScholar
2024

Consistent Diffusion Meets Tweedie: Training Exact Ambient Diffusion Models with Noisy Data

ICML 2024poster

Ambient diffusion is a recently proposed framework for training diffusion models using corrupted data. Both Ambient Diffusion and alternative SURE-based approaches for learning diffusion models from corrupted data resort to approximations which deteriorate performance. We present the first framework…

2024

DataComp-LM: In search of the next generation of training sets for language models

NeurIPS 2024poster

We introduce DataComp for Language Models, a testbed for controlled dataset experiments with the goal of improving language models. As part of DCLM, we provide a standardized corpus of 240T tokens extracted from Common Crawl, effective pretraining recipes based on the OpenLM framework, and a broad s…

Cited by 64SourcePDFScholar
2024

Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models

NeurIPS 2024poster

Using image models naively for solving inverse video problems often suffers from flickering, texture-sticking, and temporal inconsistency in generated videos. To tackle these problems, in this paper, we view frames as continuous functions in the 2D space, and videos as a sequence of continuous warpi…

2023

Ambient Diffusion: Learning Clean Distributions from Corrupted Data

NeurIPS 2023poster

We present the first diffusion-based framework that can learn an unknown distribution using only highly-corrupted samples. This problem arises in scientific applications where access to uncorrupted samples is impossible or expensive to acquire. Another benefit of our approach is the ability to train…

2023

Consistent Diffusion Models: Mitigating Sampling Drift by Learning to be Consistent

NeurIPS 2023poster

Imperfect score-matching leads to a shift between the training and the sampling distribution of diffusion models. Due to the recursive nature of the generation process, errors in previous steps yield sampling iterates that drift away from the training distribution. However, the standard training obj…

2023

DataComp: In search of the next generation of multimodal datasets

NeurIPS 2023oral

Multimodal datasets are a critical component in recent breakthroughs such as CLIP, Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms. To address this shortcoming in the machine learning ecosystem, we introduce Data…

2023

Restoration-Degradation Beyond Linear Diffusions: A Non-Asymptotic Analysis For DDIM-type Samplers

ICML 2023poster

We develop a framework for non-asymptotic analysis of deterministic samplers used for diffusion generative modeling. Several recent works have analyzed stochastic samplers using tools like Girsanov's theorem and a chain rule variant of the interpolation argument. Unfortunately, these techniques give…

Cited by 78SourcePDFScholar
2023

Solving Linear Inverse Problems Provably via Posterior Sampling with Latent Diffusion Models

NeurIPS 2023poster

We present the first framework to solve linear inverse problems leveraging pre-trained \textit{latent} diffusion models. Previously proposed algorithms (such as DPS and DDRM) only apply to \textit{pixel-space} diffusion models. We theoretically analyze our algorithm showing provable sample recover…

2022

Multitasking Models are Robust to Structural Failure: A Neural Model for Bilingual Cognitive Reserve

NeurIPS 2022accept

We find a surprising connection between multitask learning and robustness to neuron failures. Our experiments show that bilingual language models retain higher performance under various neuron perturbations, such as random deletions, magnitude pruning and weight noise. Our study is motivated by rese…

2022

Score-Guided Intermediate Level Optimization: Fast Langevin Mixing for Inverse Problems

ICML 2022spotlight

We prove fast mixing and characterize the stationary distribution of the Langevin Algorithm for inverting random weighted DNN generators. This result extends the work of Hand and Voroninski from efficient inversion to efficient posterior sampling. In practice, to allow for increased expressivity, we…

Cited by 26SourcePDFScholar
2021

Intermediate Layer Optimization for Inverse Problems using Deep Generative Models

ICML 2021spotlight

We propose Intermediate Layer Optimization (ILO), a novel optimization algorithm for solving inverse problems with deep generative models. Instead of optimizing only over the initial latent code, we progressively change the input layer obtaining successively more expressive generators. To explore th…

2021

Robust Compressed Sensing MRI with Deep Generative Priors

NeurIPS 2021poster

The CSGM framework (Bora-Jalal-Price-Dimakis'17) has shown that deep generative priors can be powerful tools for solving inverse problems. However, to date this framework has been empirically successful only on certain datasets (for example, human faces and MNIST digits), and it is known to perform…

2020

SMYRF - Efficient Attention using Asymmetric Clustering

NeurIPS 2020poster

We propose a novel type of balanced clustering algorithm to approximate attention. Attention complexity is reduced from $O(N^2)$ to $O(N \log N)$, where N is the sequence length. Our algorithm, SMYRF, uses Locality Sensitive Hashing (LSH) in a novel way by defining new Asymmetric transformations and…

2020

Your Local GAN: Designing Two Dimensional Local Attention Mechanisms for Generative Models

CVPR 2020poster

We introduce a new local sparse attention layer that preserves two-dimensional geometry and locality. We show that by just replacing the dense attention layer of SAGAN with our construction, we obtain very significant FID, Inception score and pure visual improvements. FID score is improved from 18.6…

Cited by 84PDFcodeScholar