← Search

Ivan J. Tashev

11 accepted papers

2020

Predicting Word Error Rate for Reverberant Speech

ICASSP 2020accepted

Reverberation negatively impacts the performance of automatic speech recognition (ASR). Prior work on quantifying the effect of reverberation has shown that clarity (C50), a parameter that can be estimated from the acoustic impulse response, is correlated with ASR performance. In this paper we propo…

Cited by 0SourceScholar
2019

A Sparsity Measure for Echo Density Growth in General Environments

ICASSP 2019accepted

We study the detailed temporal evolution of echo density in impulse responses for applications in acoustic analysis and rendering on general environments. For this purpose, we propose a smooth sorted density measure that yields an intuitive trend of echo density growth with time. This is fitted with…

Cited by 0SourceScholar
2019

Blind Room Volume Estimation from Single-channel Noisy Speech

ICASSP 2019accepted

Recent work on acoustic parameter estimation indicates that geometric room volume can be useful for modeling the character of an acoustic environment. However, estimating volume from audio signals remains a challenging problem. Here we propose using a convolutional neural network model to estimate t…

Cited by 0SourceScholar
2019

Improving Binaural Ambisonics Decoding by Spherical Harmonics Domain Tapering and Coloration Compensation

ICASSP 2019accepted

A powerful and flexible approach to record or encode a spatial sound scene is through spherical harmonics (SHs), or Ambisonics. An SH-encoded scene can be rendered binaurally by applying SH-encoded head-related transfer functions (HRTFs). Limitations of the recording equipment or computational const…

Cited by 0SourceScholar
2018

Spatial Audio Feature Discovery with Convolutional Neural Networks

ICASSP 2018accepted

The advent of mixed reality consumer products brings about a pressing need to develop and improve spatial sound rendering techniques for a broad user base. Despite a large body of prior work, the precise nature and importance of various sound localization cues and how they should be personalized for…

Cited by 33SourceScholar
2016

Applications of 3D spherical transforms to personalization of head-related transfer functions

ICASSP 2016accepted

Head-related transfer functions (HRTFs) depend on the shape of the human head and ears, motivating HRTF personalization methods that detect and exploit morphological similarities between subjects in an HRTF database and a new user. Prior work determined similarity from sets of morphological paramete…

Cited by 10SourceScholar
2016

BFGUI: An interactive tool for the synthesis and analysis of microphone array beamformers

ICASSP 2016accepted

Microphone arrays are beneficial for distant speech capture because the signals they capture can be exploited with beamforming to suppress noise and reverberation. The theory for the design and analysis of microphone arrays is well established, however the performance of a microphone array beamforme…

Cited by 0SourceScholar
2015

Blur kernel estimation approach to blind reverberation time estimation

ICASSP 2015accepted

Reverberation time is an important parameter for characterizing acoustic environments. It is useful in many applications including acoustic scene analysis, robust automatic speech recognition and dereverberation. Given knowledge of the acoustic impulse response, reverberation time can be measured us…

Cited by 0SourceScholar
2015

Dereverberation sweet spot dilation with combined channel equalization and beamforming

ICASSP 2015accepted

Beamforming and channel equalizers can be formulated as optimal multichannel filter-and-sum operations with different objective criteria. It has been shown in previous studies that the combination of both concepts under a common framework can yield results that combine both the spatial robustness of…

Cited by 0SourceScholar
2015

Estimation of multipath propagation delays and interaural time differences from 3-D head scans

ICASSP 2015accepted

The estimation of acoustic propagation delays from a sound source to a listener's ear entrances is useful for understanding and visualising the wave propagation along the surface of the head, and necessary for individualised spatial sound rendering. The interaural time difference (ITD) is of particu…

Cited by 0SourceScholar