← Search

Ivan Tashev

13 accepted papers

2025

Functional Near-Infrared Spectroscopy Feature Extraction with Application in Workload Estimation

ICASSP 2025accepted

Functional near-infrared spectroscopy (fNIRS) is a brain imaging technique used to estimate neuronal activity by measuring blood oxygenation. In this paper, we develop and evaluate an extensive set of fNIRS features for workload estimation, combining them with respiration and heartbeat signals. Our…

Cited by 0SourceScholar
2021

Decoding Music Attention from "EEG Headphones": A User-Friendly Auditory Brain-Computer Interface

ICASSP 2021accepted

People enjoy listening to music as part of their life. This makes music an excellent choice for designing a user-friendly brain-computer interface (BCI) for long-term use. We propose a novel BCI system using music stimuli that relies on brain signals collected via Smartfones, an EEG recording device…

Cited by 0SourceScholar
2021

Towards Efficient Models for Real-Time Deep Noise Suppression

ICASSP 2021accepted

With recent research advancements, deep learning models are be-coming attractive and powerful choices for speech enhancement in real-time applications. While state-of-the-art models can achieve outstanding results in terms of speech quality and background noise reduction, the main challenge is to ob…

Cited by 0SourceScholar
2020

Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech Enhancement

ICASSP 2020accepted

This paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channel speech enhancement. Specifically, we focus on a RNN that enhances short-time speech spectra on a single-frame-in, sing…

Cited by 0SourceScholar
2019

Non-intrusive Speech Quality Assessment Using Neural Networks

ICASSP 2019accepted

Estimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this work, we present an investigation of the applicability of neura…

Cited by 0SourceScholar
2018

A Hybrid Approach to Combining Conventional and Deep Learning Techniques for Single-Channel Speech Enhancement and Recognition

ICASSP 2018accepted

Conventional speech-enhancement techniques employ statistical signal-processing algorithms. They are computationally efficient and improve speech quality even under unknown noise conditions. For these reasons, they are preferred for deployment in unpredictable environments. One limitation of these a…

Cited by 0SourceScholar
2018

Constrained Convolutional-Recurrent Networks to Improve Speech Quality with Low Impact on Recognition Accuracy

ICASSP 2018accepted

For a speech-enhancement algorithm, it is highly desirable to simultaneously improve perceptual quality and recognition rate. Thanks to computational costs and model complexities, it is challenging to train a model that effectively optimizes both metrics at the same time. In this paper, we propose a…

Cited by 0SourceScholar
2018

Limiting Numerical Precision of Neural Networks to Achieve Real-Time Voice Activity Detection

ICASSP 2018accepted

Fast and robust voice-activity detection is critical to efficiently process speech. While deep-learning based methods to detect voice have shown competitive accuracies, the best models in the literature incur over a 100 ms latency on commodity processors. Such delays are unacceptable for real-time s…

Cited by 0SourceScholar
2017

A statistical approach to semi-supervised speech enhancement with low-order non-negative matrix factorization

ICASSP 2017accepted

Compared to generic source separation, NMF for speech enhancement is relatively underexplored. When applied to the latter problem, NMF is bereft of performance consistency (across runs and data samples), esp. with small-sized dictionaries. This limitation raises the need for higher-order representat…

Cited by 0SourceScholar
2017

Learning utterance-level representations for speech emotion and age/gender recognition using deep neural networks

ICASSP 2017accepted

Accurately recognizing speaker emotion and age/gender from speech can provide better user experience for many spoken dialogue systems. In this study, we propose to use deep neural networks (DNNs) to encode each utterance into a fixed-length vector by pooling the activations of the last hidden layer…

Cited by 0SourceScholar