← Search

Grace Luo

10 accepted papers

2026

Constantly Improving Image Models Need Constantly Improving Benchmarks

ICLR 2026poster

Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture these emerging use cases, leaving a gap between community p…

Cited by 0SourcecodeScholar
2026

Learning a Generative Meta-Model of LLM Activations

ICML 2026poster

Existing approaches for manipulating neural network activations, such as PCA and SAEs, rely on strong assumptions about activation structure. We develop a generative approach that models activations with diffusion, that makes minimal assumptions and improves with data and model scale. We use this ac…

Cited by 0SourceScholar
2024

Readout Guidance: Learning Control from Diffusion Features

CVPR 2024highlight

We present Readout Guidance a method for controlling text-to-image diffusion models with learned signals. Readout Guidance uses readout heads lightweight networks trained to extract signals from the features of a pre-trained frozen diffusion model at every timestep. These readouts can encode single-…

Cited by 26SourcePDFScholar
2023

Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence

NeurIPS 2023poster

Diffusion models have been shown to be capable of generating high-quality images, suggesting that they could contain meaningful internal representations. Unfortunately, the feature maps that encode a diffusion model's internal information are spread not only over layers of the network, but also over…

2022

Focus! Relevant and Sufficient Context Selection for News Image Captioning

EMNLP 2022finding

News Image Captioning requires describing an image by leveraging additional context derived from a news article. Previous works only coarsely leverage the article to extract the necessary context, which makes it challenging for models to identify relevant events and named entities. In our paper, we…

Cited by 11SourcePDFScholar
2022

G3: Geolocation via Guidebook Grounding

EMNLP 2022finding

We demonstrate how language can improve geolocation: the task of predicting the location where an image was taken. Here we study explicit knowledge from human-written guidebooks that describe the salient and class-discriminative visual features humans use for geolocation. We propose the task of Geol…

2022

Twitter-COMMs: Detecting Climate, COVID, and Military Multimodal Misinformation

NAACL 2022long

Detecting out-of-context media, such as “miscaptioned” images on Twitter, is a relevant problem, especially in domains of high public significance. In this work we aim to develop defenses against such misinformation for the topics of Climate Change, COVID-19, and Military Vehicles. We first present…

2021

NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal Media

EMNLP 2021main

Online misinformation is a prevalent societal issue, with adversaries relying on tools ranging from cheap fakes to sophisticated deep fakes. We are motivated by the threat scenario where an image is used out of context to support a certain narrative. While some prior datasets for detecting image-tex…