← Search

Zoe Landgraf

6 accepted papers

2025

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

CVPR 2025poster

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate this through external spatial control, increasing long-term consistency but compromi…

Cited by 2SourcePDFScholar
2024

EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars

CVPR 2024poster

Head avatars animated by visual signals have gained popularity particularly in cross-driving synthesis where the driver differs from the animated character a challenging but highly practical approach. The recently presented MegaPortraits model has demonstrated state-of-the-art results in this domain…

Cited by 26SourcePDFScholar
2024

Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs

NeurIPS 2024poster

Research in auditory, visual, and audiovisual speech recognition (ASR, VSR, and AVSR, respectively) has traditionally been conducted independently. Even recent self-supervised studies addressing two or all three tasks simultaneously tend to yield separate models, leading to disjoint inference pipeli…

2022

PINs: Progressive Implicit Networks for Multi-Scale Neural Representations

ICML 2022spotlight

Multi-layer perceptrons (MLP) have proven to be effective scene encoders when combined with higher-dimensional projections of the input, commonly referred to as positional encoding. However, scenes with a wide frequency spectrum remain a challenge: choosing high frequencies for positional encoding i…

Cited by 26SourcePDFScholar
2021

SIMstack: A Generative Shape and Instance Model for Unordered Object Stacks

ICCV 2021poster

By estimating 3D shape and instances from a single view, we can capture information about the environment quickly, without the need for comprehensive scanning and multi-view fusion. Solving this task for composite scenes (such as object stacks) is challenging: occluded areas are not only ambiguous i…

Cited by 9PDFScholar
2020

Comparing View-Based and Map-Based Semantic Labelling in Real-Time SLAM

ICRA 2020poster

Generally capable Spatial AI systems must build persistent scene representations where geometric models are combined with meaningful semantic labels. The many approaches to labelling scenes can be divided into two clear groups: view-based which estimate labels from the input view-wise data and then…

Cited by 6SourceScholar