← Search

Alexander Richard

27 accepted papers

2026

DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction

CVPR 2026

We present DuoMo, a generative method that recovers human motion in world-space coordinates from unconstrained videos with noisy or incomplete observations. Reconstructing such motion requires solving a fundamental trade-off: generalizing from diverse and noisy video inputs while maintaining global

Cited by 0SourcecodeScholar
2025

A2B: Neural Rendering of Ambisonic Recordings to Binaural

ICASSP 2025accepted

This paper introduces a novel neural network model for rendering binaural audio directly from ambisonic recordings. We optimized the model end-to-end to learn a direct mapping between ambisonic and binaural signals. Our approach eliminates traditional processing steps that were required to mitigate…

Cited by 0SourceScholar
2025

AV-Flow: Transforming Text to Audio-Visual Human-like Interactions

ICCV 2025poster

We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We demonstrate human-like speech synthesis, synchronized lip motio…

Cited by 0SourcePDFScholar
2025

BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models

ICML 2025poster

Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-qual…

Cited by 0SourcePDFScholar
2025

ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling

ICASSP 2025accepted

Neural audio codecs have been widely adopted in audio-generative tasks because their compact and discrete representations are suitable for both large-language-model-style and regression-based generative models. However, most neural codecs struggle to model out-of-domain audio, resulting in error pro…

Cited by 0SourceScholar
2025

FlowDec: A flow-based full-band general audio codec with high perceptual quality

ICLR 2025poster

We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared to the prior work ScoreDec which is based on score matching, we generalize from…

2025

REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning

CVPR 2025poster

We present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal…

Cited by 0SourcePDFScholar
2025

SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding

CVPR 2025highlight

We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the under…

Cited by 0SourcePDFScholar
2024

From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations

CVPR 2024poster

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio we output multiple possibilities of gestural motion for an individual including face body and hands. The key behind our method is…

2024

Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark

CVPR 2024highlight

We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modalities. The dataset includes high-quality and densely captured room impulse response data paired with multi-view images and precise 6DoF pose tracking data for sound emitters and listen…

Cited by 13SourcePDFScholar
2024

ScoreDec: A Phase-Preserving High-Fidelity Audio Codec with a Generalized Score-Based Diffusion Post-Filter

ICASSP 2024accepted

Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative adversarial network (GAN) training is usually required for these E2E…

Cited by 0SourceScholar
2023

Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec

ICASSP 2023accepted

A good audio codec for live applications such as telecommunication is characterized by three key properties: (1) compression, i.e. the bitrate that is required to transmit the signal should be as low as possible; (2) latency, i.e. encoding and decoding the signal needs to be fast enough to enable co…

Cited by 0SourceScholar
2023

Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural Speech

ICASSP 2023accepted

We propose NORD: a novel framework for estimating the relative depth between two binaural speech recordings. In contrast to existing depth estimation techniques, ours only requires audio signals as input. We trained the framework to solve depth preference (i.e. which input perceptually sounds closer…

Cited by 0SourceScholar
2023

Novel-View Acoustic Synthesis

CVPR 2023poster

We introduce the novel-view acoustic synthesis (NVAS) task: given the sight and sound observed at a source viewpoint, can we synthesize the sound of that scene from an unseen target viewpoint? We propose a neural rendering approach: Visually-Guided Acoustic Synthesis (ViGAS) network that learns to s…

2023

Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and Audio

NeurIPS 2023spotlight

While 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we present a model that can generate accurate 3D spatial audio for f…

2022

Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis

CVPR 2022oral

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet, state-of-the-art approaches still struggle to generate clean, realistic speech…

Cited by 45PDFcodeScholar
2022

Conditional Diffusion Probabilistic Model for Speech Enhancement

ICASSP 2022accepted

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models have shown strong potential in speech synthesis, they are still lagging behind in speech enhancement. This work leverage…

Cited by 0SourceScholar
2022

Deep Impulse Responses: Estimating and Parameterizing Filters with Deep Networks

ICASSP 2022accepted

Impulse response estimation in high noise and in-the-wild settings, with minimal control of the underlying data distributions, is a challenging problem. We propose a novel framework for parameterizing and estimating impulse responses based on recent advances in neural representation learning. Our fr…

Cited by 0SourceScholar
2022

LiP-Flow: Learning Inference-Time Priors for Codec Avatars via Normalizing Flows in Latent Space

ECCV 2022poster

"Neural face avatars that are trained from multi-view data captured in camera domes can produce photo-realistic 3D reconstructions. However, at inference time, they must be driven by limited inputs such as partial views recorded by headset-mounted cameras or a front-facing camera, and sparse facial…

Cited by 1SourcePDFScholar
2021

Implicit HRTF Modeling Using Temporal Convolutional Networks

ICASSP 2021accepted

Estimation of accurate head-related transfer functions (HRTFs) is crucial to achieve realistic binaural acoustic experiences. HRTFs depend on source/listener locations and are therefore expensive and cumbersome to measure; traditional approaches require listener-dependent measurements of HRTFs at th…

Cited by 0SourceScholar
2021

MeshTalk: 3D Face Animation From Speech Using Cross-Modality Disentanglement

ICCV 2021poster

This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible co-articulation or rely on person-specific models that limit their sca…

Cited by 241PDFcodeScholar
2021

Neural Synthesis of Binaural Speech From Mono Audio

ICLR 2021oral

We present a neural rendering approach for binaural sound synthesis that can produce realistic and spatially accurate binaural sound in realtime. The network takes, as input, a single-channel audio source and synthesizes, as output, two-channel binaural sound, conditioned on the relative position an…

Cited by 76SourcePDFScholar
2018

Action Sets: Weakly Supervised Action Segmentation Without Ordering Constraints

CVPR 2018poster

Action detection and temporal segmentation of actions in videos are topics of increasing interest. While fully supervised systems have gained much attention lately, full annotation of each action within the video is costly and impractical for large amounts of video data. Thus, weakly supervised acti…

Cited by 115SourcePDFScholar
2018

NeuralNetwork-Viterbi: A Framework for Weakly Supervised Video Learning

CVPR 2018poster

Video learning is an important task in computer vision and has experienced increasing interest over the recent years. Since even a small amount of videos easily comprises several million frames, methods that do not rely on a frame-level annotation are of special importance. In this work, we propose…

Cited by 173SourcePDFScholar
2018

When Will You Do What? - Anticipating Temporal Occurrences of Activities

CVPR 2018poster

Analyzing human actions in videos has gained increased attention recently. While most works focus on classifying and labeling observed video frames or anticipating the very recent future, making long-term predictions over more than just a few seconds is a task with many practical applications that h…

2017

Weakly Supervised Action Learning With RNN Based Fine-To-Coarse Modeling

CVPR 2017oral

We present an approach for weakly supervised learning of human actions. Given a set of videos and an ordered list of the occurring actions, the goal is to infer start and end frames of the related action classes within the video and to train the respective action classifiers without any need for han…

Cited by 264PDFcodeScholar