← Search

A. Sophia Koepke

12 accepted papers

2026

It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models

CVPR 2026

Contemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prompt. Previous work has attempted to address this issue by steering the model using guidance mechanisms, or by generating a large pool of candidates and r

Cited by 0SourcecodeScholar
2025

VGGSounder: Audio-Visual Evaluations for Foundation Models

ICCV 2025poster

Designing effective foundation models requires high-quality evaluation datasets. With the emergence of audio-visual foundation models, reliable assessment of their multi-modal understanding is essential. The current gold standard for evaluating audio-visual understanding is the popular classificatio…

Cited by 0SourcePDFScholar
2024

A Sound Approach: Using Large Language Models to Generate Audio Descriptions for Egocentric Text-Audio Retrieval

ICASSP 2024accepted

Video databases from the internet are a valuable source of text-audio retrieval datasets. However, given that sound and vision streams represent different "views" of the data, treating visual descriptions as audio descriptions is far from optimal. Even if audio class labels are present, they commonl…

Cited by 0SourceScholar
2024

Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model

ICLR 2024spotlight

Training deep networks requires various design decisions regarding for instance their architecture, data augmentation, or optimization. In this work, we find these training variations to result in networks learning unique feature sets from the data. Using public model libraries comprising thousands…

2023

Image-Free Classifier Injection for Zero-Shot Classification

ICCV 2023poster

Zero-shot learning models achieve remarkable results on image classification for samples from classes that were not seen during training. However, such models must be trained from scratch with specialised methods: therefore, access to a training dataset is required when the need for zero-shot classi…

Cited by 16PDFcodeScholar
2023

Waffling Around for Performance: Visual Classification with Random Words and Broad Concepts

ICCV 2023poster

The visual classification performance of vision-language models such as CLIP has been shown to benefit from additional semantic knowledge from large language models (LLMs) such as GPT-3. In particular, averaging over LLM-generated class descriptors, e.g. "waffle, which has a round shape", can notabl…

Cited by 86PDFcodeScholar
2022

Audio-Visual Generalised Zero-Shot Learning With Cross-Modal Attention and Language

CVPR 2022poster

Learning to classify video data from classes not included in the training data, i.e. video-based zero-shot learning, is challenging. We conjecture that the natural alignment between the audio and visual modalities in video data provides a rich training signal for learning discriminative multi-modal…

Cited by 72PDFcodeScholar
2022

PlanT: Explainable Planning Transformers via Object-Level Representations

CoRL 2022poster

Planning an optimal route in a complex environment requires efficient reasoning about the surrounding scene. While human drivers prioritize important objects and ignore details not relevant to the decision, learning-based planners typically extract features from dense, high-dimensional grid represen…

Cited by 114SourcecodeScholar
2022

Temporal and Cross-Modal Attention for Audio-Visual Zero-Shot Learning

ECCV 2022poster

"Audio-visual generalised zero-shot learning for video classification requires understanding the relations between the audio and visual information in order to be able to recognise samples from novel, previously unseen classes at test time. The natural semantic and temporal alignment between audio a…

2021

Distilling Audio-Visual Knowledge by Compositional Contrastive Learning

CVPR 2021poster

Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous modalities, even though these data modalities may not be semantically correlated.…

Cited by 95PDFcodeScholar
2020

Sight to Sound: An End-to-End Approach for Visual Piano Transcription

ICASSP 2020accepted

Automatic music transcription has primarily focused on transcribing audio to a symbolic music representation (e.g. MIDI or sheet music). However, audio-only approaches often struggle with polyphonic instruments and background noise. In contrast, visual information (e.g. a video of an instrument bein…

Cited by 0SourceScholar
2018

X2Face: A network for controlling face generation using images, audio, and pose codes

ECCV 2018poster

The objective of this paper is a neural network model that controls the pose and expression of a given face, using another face or modality (e.g. audio). This model can then be used for lightweight, sophisticated video and image editing. We make the following three contributions. First, we introduce…

Cited by 511SourcePDFScholar