← Search

Matthias Grundmann

7 accepted papers

2024

Binaural Angular Separation Network

ICASSP 2024accepted

We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omnidirectional microphones without needing to collect real RIRs. By relying…

Cited by 0SourceScholar
2024

PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models

CVPR 2024poster

Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using reinforcement learning (RL) to maximize rewards that reflect human preference. However in the vision domain existing RL…

Cited by 15SourcePDFScholar
2024

STREAMVC: Real-Time Low-Latency Voice Conversion

ICASSP 2024accepted

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces the resulting waveform at low latency from the input signal even on a mobile pl…

Cited by 0SourceScholar
2023

Guided Speech Enhancement Network

ICASSP 2023accepted

High quality speech capture has been widely studied for both voice communication and human computer interface reasons. To improve the capture performance, we can often find multi-microphone speech enhancement techniques deployed on various devices. Multi-microphone speech enhancement problem is ofte…

Cited by 0SourceScholar
2023

Semi-Implicit Denoising Diffusion Models (SIDDMs)

NeurIPS 2023poster

Despite the proliferation of generative models, achieving fast sampling during inference without compromising sample diversity and quality remains challenging. Existing models such as Denoising Diffusion Probabilistic Models (DDPM) deliver high-quality, diverse samples but are slowed by an inherentl…

2023

Towards Authentic Face Restoration with Iterative Diffusion Models and Beyond

ICCV 2023poster

An authentic face restoration system is becoming increasingly demanding in many computer vision applications, e.g., image enhancement, video communication, and taking portrait. Most of the advanced face restoration models can recover high-quality faces from low-quality ones but usually fail to faith…

Cited by 17PDFcodeScholar
2021

Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild With Pose Annotations

CVPR 2021poster

3D object detection has recently become popular due to many applications in robotics, augmented reality, autonomy, and image retrieval. We introduce the Objectron dataset to advance the state of the art in 3D object detection and foster new research and applications, such as 3D object tracking, view…

Cited by 231PDFcodeScholar