← Search

Cory Stephenson

6 accepted papers

2026

What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment

AAAI 2026technical

Automated evaluation of generative text-to-image models remains a challenging problem. Recent works have proposed using multimodal LLMs to judge the quality of images, but these works offer little insight into how multimodal LLMs make use of concepts relevant to humans, such as image style or compos

Cited by 0SourcePDFScholar
2024

CommonCanvas: Open Diffusion Models Trained on Creative-Commons Images

CVPR 2024poster

We train a set of open text-to-image (T2I) diffusion models on a dataset of curated Creative-Commons-licensed (CC) images which yields models that are competitive with Stable Diffusion 2 (SD2). This task presents two challenges: (1) high-resolution CC images lack the captions necessary to train T2I…

Cited by 30SourcePDFScholar
2021

On the geometry of generalization and memorization in deep neural networks

ICLR 2021poster

Understanding how large neural networks avoid memorizing training data is key to explaining their high generalization performance. To examine the structure of when and where memorization occurs in a deep network, we use a recently developed replica-based mean field theoretic geometric analysis metho…

Cited by 87SourcePDFScholar
2020

Emergence of Separable Manifolds in Deep Language Representations

ICML 2020poster

Deep neural networks (DNNs) have shown much empirical success in solving perceptual tasks across various cognitive modalities. While they are only loosely inspired by the biological brain, recent studies report considerable similarities between representations extracted from task-optimized DNNs and…

2019

Adversarially Trained Autoencoders for Parallel-data-free Voice Conversion

ICASSP 2019accepted

We present a method for converting the voices between a set of speakers. Our method is based on training multiple autoencoder paths, where there is a single speaker-independent encoder and multiple speaker-dependent decoders. The autoencoders are trained with an addition of an adversarial loss which…

Cited by 0SourceScholar
2019

Untangling in Invariant Speech Recognition

NeurIPS 2019poster

Encouraged by the success of deep convolutional neural networks on a variety of visual tasks, much theoretical and experimental work has been aimed at understanding and interpreting how vision networks operate. At the same time, deep neural networks have also achieved impressive performance in audi…