← Search

Olivia Wiles

13 accepted papers

2026

Dynamic Classifier-Free Diffusion Guidance via Online Feedback

ICLR 2026poster

Classifier-free guidance (CFG) is a cornerstone of text-to-image diffusion models, yet its effectiveness is limited by the use of static guidance scales. This ``one-size-fits-all'' approach fails to adapt to the diverse requirements of different prompts; moreover, prior solutions like gradient-based…

Cited by 0SourceScholar
2025

Century: A Framework and Dataset for Evaluating Historical Contextualisation of Sensitive Images

ICLR 2025spotlight

How do multi-modal generative models describe images of recent historical events and figures, whose legacies may be nuanced, multifaceted, or contested? This task necessitates not only accurate visual recognition, but also socio-cultural knowledge and cross-modal reasoning. To address this evaluati…

Cited by 0SourcePDFScholar
2025

Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human rating

ICLR 2025spotlight

While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While many metrics and benchmarks have been proposed to evaluate T2I models and alignment metrics, the impact of the evaluation components (prompt sets, human…

Cited by 12SourcePDFScholar
2024

Evaluating Model Bias Requires Characterizing its Mistakes

ICML 2024poster

The ability to properly benchmark model performance in the face of spurious correlations is important to both build better predictors and increase confidence that models are operating as intended. We demonstrate that characterizing (as opposed to simply quantifying) model mistakes across subgroups i…

Cited by 2SourcePDFScholar
2024

Evaluating Numerical Reasoning in Text-to-Image Models

NeurIPS 2024poster

Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensively evaluate a range of text-to-image models on numerical reasoning tasks of varying difficulty, and show that even the mo…

2022

A Fine-Grained Analysis on Distribution Shift

ICLR 2022oral

Robustness to distribution shifts is critical for deploying machine learning models in the real world. Despite this necessity, there has been little work in defining the underlying mechanisms that cause these shifts and evaluating the robustness of algorithms across multiple, different distribution…

2022

Defending Against Image Corruptions Through Adversarial Augmentations

ICLR 2022poster

Modern neural networks excel at image classification, yet they remain vulnerable to common image corruptions such as blur, speckle noise or fog. Recent methods that focus on this problem, such as AugMix and DeepAugment, introduce defenses that operate in expectation over a distribution of image corr…

Cited by 54SourcePDFScholar
2021

Data Augmentation Can Improve Robustness

NeurIPS 2021poster

Adversarial training suffers from robust overfitting, a phenomenon where the robust test accuracy starts to decrease during training. In this paper, we focus on reducing robust overfitting by using common data augmentation schemes. We demonstrate that, contrary to previous findings, when combined wi…

2021

Improving Robustness using Generated Data

NeurIPS 2021poster

Recent work argues that robust training requires substantially larger datasets than those required for standard classification. On CIFAR-10 and CIFAR-100, this translates into a sizable robust-accuracy gap between models trained solely on data from the original training set and those trained with ad…

Cited by 353SourcePDFScholar
2020

Sight to Sound: An End-to-End Approach for Visual Piano Transcription

ICASSP 2020accepted

Automatic music transcription has primarily focused on transcribing audio to a symbolic music representation (e.g. MIDI or sheet music). However, audio-only approaches often struggle with polyphonic instruments and background noise. In contrast, visual information (e.g. a video of an instrument bein…

Cited by 0SourceScholar
2020

SynSin: End-to-End View Synthesis From a Single Image

CVPR 2020oral

View synthesis allows for the generation of new views of a scene given one or more images. This is challenging; it requires comprehensively understanding the 3D scene from images. As a result, current methods typically use multiple images, train on ground-truth depth, or are limited to synthetic dat…

Cited by 510PDFcodeScholar
2018

X2Face: A network for controlling face generation using images, audio, and pose codes

ECCV 2018poster

The objective of this paper is a neural network model that controls the pose and expression of a given face, using another face or modality (e.g. audio). This model can then be used for lightweight, sophisticated video and image editing. We make the following three contributions. First, we introduce…

Cited by 511SourcePDFScholar