← Search

Dejan Markovic

13 accepted papers

2025

A2B: Neural Rendering of Ambisonic Recordings to Binaural

ICASSP 2025accepted

This paper introduces a novel neural network model for rendering binaural audio directly from ambisonic recordings. We optimized the model end-to-end to learn a direct mapping between ambisonic and binaural signals. Our approach eliminates traditional processing steps that were required to mitigate…

Cited by 0SourceScholar
2025

BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models

ICML 2025poster

Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-qual…

Cited by 0SourcePDFScholar
2025

ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling

ICASSP 2025accepted

Neural audio codecs have been widely adopted in audio-generative tasks because their compact and discrete representations are suitable for both large-language-model-style and regression-based generative models. However, most neural codecs struggle to model out-of-domain audio, resulting in error pro…

Cited by 0SourceScholar
2025

SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding

CVPR 2025highlight

We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the under…

Cited by 0SourcePDFScholar
2024

ScoreDec: A Phase-Preserving High-Fidelity Audio Codec with a Generalized Score-Based Diffusion Post-Filter

ICASSP 2024accepted

Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative adversarial network (GAN) training is usually required for these E2E…

Cited by 0SourceScholar
2023

Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec

ICASSP 2023accepted

A good audio codec for live applications such as telecommunication is characterized by three key properties: (1) compression, i.e. the bitrate that is required to transmit the signal should be as low as possible; (2) latency, i.e. encoding and decoding the signal needs to be fast enough to enable co…

Cited by 0SourceScholar
2023

Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural Speech

ICASSP 2023accepted

We propose NORD: a novel framework for estimating the relative depth between two binaural speech recordings. In contrast to existing depth estimation techniques, ours only requires audio signals as input. We trained the framework to solve depth preference (i.e. which input perceptually sounds closer…

Cited by 0SourceScholar
2023

Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and Audio

NeurIPS 2023spotlight

While 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we present a model that can generate accurate 3D spatial audio for f…

2021

Implicit HRTF Modeling Using Temporal Convolutional Networks

ICASSP 2021accepted

Estimation of accurate head-related transfer functions (HRTFs) is crucial to achieve realistic binaural acoustic experiences. HRTFs depend on source/listener locations and are therefore expensive and cumbersome to measure; traditional approaches require listener-dependent measurements of HRTFs at th…

Cited by 0SourceScholar
2021

Neural Synthesis of Binaural Speech From Mono Audio

ICLR 2021oral

We present a neural rendering approach for binaural sound synthesis that can produce realistic and spatially accurate binaural sound in realtime. The network takes, as input, a single-channel audio source and synthesizes, as output, two-channel binaural sound, conditioned on the relative position an…

Cited by 76SourcePDFScholar
2019

Soundfield Reconstruction in Reverberant Environments Using Higher-order Microphones and Impulse Response Measurements

ICASSP 2019accepted

This paper addresses the problem of soundfield reconstruction over a large area using a distributed array of higher-order microphones. Given an area enclosed by the array, one can distinguish between two components of the soundfield: the interior soundfield generated by sources outside of the enclos…

Cited by 0SourceScholar
2018

A Weighted Least Squares Beam Shaping Technique for Sound Field Control

ICASSP 2018accepted

A weighted least squares beam shaping technique for sound field control using a loudspeaker array is proposed. Given a desired spatial response at prescribed control points, the space-time filter is designed by solving a least squares minimization problem. To reduce the computational effort, we prop…

Cited by 0SourceScholar
2016

A linear operator for the computation of soundfield maps

ICASSP 2016accepted

In the process of soundfield imaging, as defined in the literature, a microphone array is subdivided into overlapping sub-arrays and soundfield images are obtained by juxtaposition of spatial spectra computed from individual subarray data. In this paper we show that the whole process can be convenie…

Cited by 0SourceScholar