← Search

Israel D. Gebru

10 accepted papers

2025

A2B: Neural Rendering of Ambisonic Recordings to Binaural

ICASSP 2025accepted

This paper introduces a novel neural network model for rendering binaural audio directly from ambisonic recordings. We optimized the model end-to-end to learn a direct mapping between ambisonic and binaural signals. Our approach eliminates traditional processing steps that were required to mitigate…

Cited by 0SourceScholar
2025

BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models

ICML 2025poster

Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-qual…

Cited by 0SourcePDFScholar
2025

ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling

ICASSP 2025accepted

Neural audio codecs have been widely adopted in audio-generative tasks because their compact and discrete representations are suitable for both large-language-model-style and regression-based generative models. However, most neural codecs struggle to model out-of-domain audio, resulting in error pro…

Cited by 0SourceScholar
2025

SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding

CVPR 2025highlight

We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the under…

Cited by 0SourcePDFScholar
2024

Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark

CVPR 2024highlight

We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modalities. The dataset includes high-quality and densely captured room impulse response data paired with multi-view images and precise 6DoF pose tracking data for sound emitters and listen…

Cited by 13SourcePDFScholar
2024

ScoreDec: A Phase-Preserving High-Fidelity Audio Codec with a Generalized Score-Based Diffusion Post-Filter

ICASSP 2024accepted

Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative adversarial network (GAN) training is usually required for these E2E…

Cited by 0SourceScholar
2023

Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec

ICASSP 2023accepted

A good audio codec for live applications such as telecommunication is characterized by three key properties: (1) compression, i.e. the bitrate that is required to transmit the signal should be as low as possible; (2) latency, i.e. encoding and decoding the signal needs to be fast enough to enable co…

Cited by 0SourceScholar
2023

Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural Speech

ICASSP 2023accepted

We propose NORD: a novel framework for estimating the relative depth between two binaural speech recordings. In contrast to existing depth estimation techniques, ours only requires audio signals as input. We trained the framework to solve depth preference (i.e. which input perceptually sounds closer…

Cited by 0SourceScholar
2021

Implicit HRTF Modeling Using Temporal Convolutional Networks

ICASSP 2021accepted

Estimation of accurate head-related transfer functions (HRTFs) is crucial to achieve realistic binaural acoustic experiences. HRTFs depend on source/listener locations and are therefore expensive and cumbersome to measure; traditional approaches require listener-dependent measurements of HRTFs at th…

Cited by 0SourceScholar
2021

Neural Synthesis of Binaural Speech From Mono Audio

ICLR 2021oral

We present a neural rendering approach for binaural sound synthesis that can produce realistic and spatially accurate binaural sound in realtime. The network takes, as input, a single-channel audio source and synthesizes, as output, two-channel binaural sound, conditioned on the relative position an…

Cited by 76SourcePDFScholar