← Search

Marc Delcroix

72 accepted papers

2026

FLEXIO: FLEXIBLE SINGLE- AND MULTI-CHANNEL SPEECH SEPARATION AND ENHANCEMENT

ICASSP 2026oral

Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speak…

Cited by 0SourcePDFScholar
2026

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

ICML 2026poster

Robust classification in noisy environments remains a fundamental challenge in machine learning. Standard approaches typically treat signal enhancement and classification as separate, sequential stages: first enhancing the signal and then applying a classifier. This approach fails to leverage the se…

Cited by 0SourceScholar
2026

LOOSE COUPLING OF SPECTRAL AND SPATIAL MODELS FOR MULTI-CHANNEL DIARIZATION AND ENHANCEMENT OF MEETINGS IN DYNAMIC ENVIRONMENTS

ICASSP 2026poster

Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in meeting transcription. However, there is no one-to-one mapping of positions in space to speakers if speakers move. Here, we…

Cited by 0SourcePDFScholar
2026

SPATIALLY AWARE SELF-SUPERVISED MODELS FOR MULTI-CHANNEL NEURAL SPEAKER DIARIZATION

ICASSP 2026poster

Self-supervised models such as WavLM have demonstrated strong performance for neural speaker diarization. However, these models are typically pre-trained on single-channel recordings, limiting their effectiveness in multi-channel scenarios. Existing diarization systems built on these models often re…

Cited by 0SourcePDFScholar
2025

A Hybrid Probabilistic-Deterministic Model Recursively Enhancing Speech

ICASSP 2025accepted

This paper introduces Probabilistic-Deterministic Recursive Enhancement (PDRE), an innovative iterative Speech Enhancement (SE) approach that integrates probabilistic and deterministic methodologies. Recent advancements in diffusion models have demonstrated the exceptional effectiveness of probabili…

Cited by 0SourceScholar
2025

Alignment-Free Training for Transducer-based Multi-Talker ASR

ICASSP 2025accepted

Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to achieve recognition without relying on costly front-end source separation. MT-RNNT is conventionally implemented using arch…

Cited by 0SourceScholar
2025

Bridging Speech and Text Foundation Models with ReShape Attention

ICASSP 2025accepted

This paper investigates cascade approaches bridging speech and text foundation models (FMs) for speech translation (ST). We address the limitations of cascade systems which suffer from the propagation of speech recognition errors and the lack of access to acoustic information. We propose a ReShape A…

Cited by 0SourceScholar
2025

Guided Speaker Embedding

ICASSP 2025accepted

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment…

Cited by 0SourceScholar
2025

Mamba-based Segmentation Model for Speaker Diarization

ICASSP 2025accepted

Mamba is a newly proposed architecture that behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are…

Cited by 0SourceScholar
2025

Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation

ICASSP 2025accepted

This paper proposes a speaker counting scheme using multichannel microphones for end-to-end neural diarization with a vector clustering (EEND-VC) speaker diarization pipeline. The EEND-VC-based system estimates the number of speakers by clustering speaker embeddings from small chunks. However, conve…

Cited by 0SourceScholar
2025

SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model

ICASSP 2025accepted

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the sa…

Cited by 0SourceScholar
2025

TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models

ICASSP 2025accepted

Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in nois…

Cited by 0SourceScholar
2024

Discriminative Training of VBx Diarization

ICASSP 2024accepted

Bayesian HMM clustering of x-vector sequences (VBx) has become a widely adopted diarization baseline model in publications and challenges. It uses an HMM to model speaker turns, a generatively trained probabilistic linear discriminant analysis (PLDA) for speaker distribution modeling, and Bayesian i…

Cited by 0SourceScholar
2024

How Does End-To-End Speech Recognition Training Impact Speech Enhancement Artifacts?

ICASSP 2024accepted

Jointly training a speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end has been investigated as a way to mitigate the influence of processing distortion generated by single-channel SE on ASR. In this paper, we investigate the effect of such joint training on the sign…

Cited by 0SourceScholar
2024

NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization

ICASSP 2024accepted

This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)based dereverberation as a front end, and separately applies end-to-end neural diarization with vector clustering…

Cited by 0SourceScholar
2024

Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task Loss

ICASSP 2024accepted

Array processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone locati…

Cited by 0SourceScholar
2024

Noise-Robust Zero-Shot Text-to-Speech Synthesis Conditioned on Self-Supervised Speech-Representation Model with Adapters

ICASSP 2024accepted

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this approach suffers from degradation in speech synthesis quality when…

Cited by 0SourceScholar
2024

Online Target Sound Extraction with Knowledge Distillation from Partially Non-Causal Teacher

ICASSP 2024accepted

Target Sound Extraction (TSE) is a technique for extracting sound events belonging to a target sound class in a mixture using a Deep Neural Network (DNN). Offline TSE that uses non-causal models has achieved high extraction performance. However, many applications require online processing. Simply co…

Cited by 0SourceScholar
2024

Target Speech Extraction with Pre-Trained Self-Supervised Learning Models

ICASSP 2024accepted

Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been fully exploited. TSE aims to extract the speech of a target speaker in a mixture guided by enrollment utterances. We exp…

Cited by 0SourceScholar
2024

Train Long and Test Long: Leveraging Full Document Contexts in Speech Processing

ICASSP 2024accepted

The quadratic memory complexity of self-attention has generally restricted Transformer-based models to utterance-based speech processing, preventing models from leveraging long-form contexts. A common solution has been to formulate long-form speech processing into a streaming problem, only using lim…

Cited by 0SourceScholar
2024

What Do Self-Supervised Speech and Speaker Models Learn? New Findings from a Cross Model Layer-Wise Analysis

ICASSP 2024accepted

Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representations. In contrast, speaker SSL models, exemplified by DINO-based models, adopt utt…

Cited by 0SourceScholar
2023

Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition

ICASSP 2023accepted

We propose a new shallow fusion (SF) method to exploit an external backward language model (BLM) for end-to-end automatic speech recognition (ASR). The BLM has complementary characteristics with a forward language model (FLM), and the effectiveness of their combination has been confirmed by rescorin…

Cited by 0SourceScholar
2023

Leveraging Large Text Corpora For End-To-End Speech Summarization

ICASSP 2023accepted

End-to-end speech summarization (E2E SSum) is a technique to directly generate summary sentences from speech. Compared with the cascade approach, which combines automatic speech recognition (ASR) and text summarization models, the E2E approach is more promising because it mitigates ASR errors, incor…

Cited by 0SourceScholar
2023

On Word Error Rate Definitions and Their Efficient Computation for Multi-Speaker Speech Recognition Systems

ICASSP 2023accepted

We propose a general framework to compute the word error rate (WER) of ASR systems that process recordings containing multiple speakers at their input and that produce multiple output word sequences (MIMO). Such ASR systems are typically required, e.g., for meeting transcription. We provide an effic…

Cited by 0SourceScholar
2023

Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text Encoders

ICASSP 2023accepted

Speech summarization requires processing several minute-long speech sequences to allow exploiting the whole context of a spoken document. A conventional approach is a cascade of automatic speech recognition (ASR) and text summarization (TS). However, the cascade systems are sensitive to ASR errors.…

Cited by 0SourceScholar
2022

Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration

ICASSP 2022accepted

In this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (…

Cited by 0SourceScholar
2022

Integrating Multiple ASR Systems into NLP Backend with Attention Fusion

ICASSP 2022accepted

Spoken language processing (SLP) systems such as speech summarization and translation can be achieved by cascade models. It combines an automatic speech recognition (ASR) frontend and a natural language processing (NLP) backend including machine translation (MT) or text summarization (TS). With this…

Cited by 0SourceScholar
2022

Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models

ICASSP 2022accepted

We investigate the effectiveness of using a large ensemble of advanced neural language models (NLMs) for lattice rescoring on automatic speech recognition (ASR) hypotheses. Previous studies have reported the effectiveness of combining a small number of NLMs. In contrast, in this study, we combine up…

Cited by 0SourceScholar
2022

Learning to Enhance or Not: Neural Network-Based Switching of Enhanced and Observed Signals for Overlapping Speech Recognition

ICASSP 2022accepted

The combination of a deep neural network (DNN) -based speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end is a widely used approach to implement overlapping speech recognition. However, the SE front-end generates processing artifacts that can degrade the ASR performa…

Cited by 0SourceScholar
2022

SA-SDR: A Novel Loss Function for Separation of Meeting Style Data

ICASSP 2022accepted

Many state-of-the-art neural network-based source separation systems use the averaged Signal-to-Distortion Ratio (SDR) as a training objective function. The basic SDR is, however, undefined if the network reconstructs the reference signal perfectly or if the reference signal contains silence, e.g.,…

Cited by 0SourceScholar
2022

Tight Integration Of Neural- And Clustering-Based Diarization Through Deep Unfolding Of Infinite Gaussian Mixture Model

ICASSP 2022accepted

Speaker diarization has been investigated extensively as an important central task for meeting analysis. Recent trend shows that integration of end-to-end neural (EEND)- and clustering-based diarization is a promising approach to handle realistic conversational data containing overlapped speech with…

Cited by 0SourceScholar
2021

BLSTM-Based Confidence Estimation for End-to-End Speech Recognition

ICASSP 2021accepted

Confidence estimation, in which we estimate the reliability of each recognized token (e.g., word, sub-word, and character) in automatic speech recognition (ASR) hypotheses and detect incorrectly recognized tokens, is an important function for developing ASR applications. In this study, we perform co…

Cited by 0SourceScholar
2021

Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech Separation

ICASSP 2021accepted

Time-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech enhancement and source separation. Here, we propose to combine ne…

Cited by 0SourceScholar
2021

Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain

ICASSP 2021accepted

Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning unique speaker identities. Recently, several approaches utili…

Cited by 0SourceScholar
2021

Dual-Path Modeling for Long Recording Speech Separation in Meetings

ICASSP 2021accepted

The continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straightforward extension of conventional utterance-level speech separation to the CSS task is to segment the long recording wit…

Cited by 0SourceScholar
2021

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

ICASSP 2021accepted

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both singlechannel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performan…

Cited by 0SourceScholar
2021

Integrating End-to-End Neural and Clustering-Based Diarization: Getting the Best of Both Worlds

ICASSP 2021accepted

Recent diarization technologies can be categorized into two approaches, i.e., clustering and end-to-end neural approaches, which have different pros and cons. The clustering-based approaches assign speaker labels to speech regions by clustering speaker embeddings such as x-vectors. While it can be s…

Cited by 0SourceScholar
2021

Neural Network-Based Virtual Microphone Estimator

ICASSP 2021accepted

Developing microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g., based on several physical model assumptions. However, such…

Cited by 0SourceScholar
2021

Speaker Activity Driven Neural Speech Extraction

ICASSP 2021accepted

Target speech extraction, which extracts the speech of a target speaker in a mixture given auxiliary speaker clues, has recently received increased interest. Various clues have been investigated such as pre-recorded enrollment utterances, direction information, or video of the target speaker. In thi…

Cited by 0SourceScholar
2020

A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker Tracking

ICASSP 2020accepted

Audiovisual speaker tracking is an application that has been tackled by a wide range of classical approaches based on Gaussian filters, most notably the well-known Kalman filter. Recently, a specific Kalman filter implementation was proposed for this task, which incorporated dynamic stream weights t…

Cited by 0SourceScholar
2020

Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain Beamformer

ICASSP 2020accepted

Recent studies have shown that acoustic beamforming using a microphone array plays an important role in the construction of high-performance automatic speech recognition (ASR) systems, especially for noisy and overlapping speech conditions. In parallel with the success of multichannel beamforming fo…

Cited by 0SourceScholar
2020

DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source Separation

ICASSP 2020accepted

In this article, we investigate an integrated mask-based convolutional beamforming method for performing simultaneous denoising, dereverberation, and source separation. Conventionally, it is difficult for neural network-supported mask-based source separation to perform denoising and dereverberation…

Cited by 0SourceScholar
2020

End-to-End Training of Time Domain Audio Separation and Recognition

ICASSP 2020accepted

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recogniti…

Cited by 0SourceScholar
2020

Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances

ICASSP 2020accepted

This paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and tes…

Cited by 0SourceScholar
2020

Improving Noise Robust Automatic Speech Recognition with Single-Channel Time-Domain Enhancement Network

ICASSP 2020accepted

With the advent of deep learning, research on noise-robust automatic speech recognition (ASR) has progressed rapidly. However, ASR performance in noisy conditions of single-channel systems remains unsatisfactory. Indeed, most single-channel speech enhancement (SE) methods (denoising) have brought on…

Cited by 0SourceScholar
2020

Improving Speaker Discrimination of Target Speech Extraction With Time-Domain Speakerbeam

ICASSP 2020accepted

Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation utterance of the target speaker to extract his/her voice characteristics that are…

Cited by 0SourceScholar
2020

Speech Enhancement Using Self-Adaptation and Multi-Head Self-Attention

ICASSP 2020accepted

This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studies of deep neural network (DNN)-based speech enhancement mainly focus on buildin…

Cited by 0SourceScholar
2020

Tackling Real Noisy Reverberant Meetings with All-Neural Source Separation, Counting, and Diarization System

ICASSP 2020accepted

Automatic meeting analysis is an essential fundamental technology required to let, e.g. smart devices follow and respond to our conversations. To achieve an optimal automatic meeting analysis, we previously proposed an all-neural approach that jointly solves source separation, speaker diarization an…

Cited by 0SourceScholar
2019

A Unified Framework for Feature-based Domain Adaptation of Neural Network Language Models

ICASSP 2019accepted

An important task for language models is the adaptation of general-domain models to specific target domains. For neural network-based language models, feature-based domain adaptation has been a popular method in previous research. Conventional methods use an adaptation feature providing context info…

Cited by 0SourceScholar
2019

A Unified Framework for Neural Speech Separation and Extraction

ICASSP 2019accepted

The development of deep learning techniques has triggered the active investigation of neural network-based speech enhancement approaches. In particular, single-channel blind (uninformed) speech separation and speaker-aware (informed) speech extraction have received increased interest. Blind speech s…

Cited by 0SourceScholar
2019

All-neural Online Source Separation, Counting, and Diarization for Meeting Analysis

ICASSP 2019accepted

Automatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant…

Cited by 0SourceScholar
2019

Compact Network for Speakerbeam Target Speaker Extraction

ICASSP 2019accepted

Speech separation that separates a mixture of speech signals into each of its sources has been an active research topic for a long time and has seen recent progress with the advent of deep learning. A related problem is target speaker extraction, i.e. extraction of only speech of a target speaker ou…

Cited by 0SourceScholar
2019

Estimation of Sampling Frequency Mismatch between Distributed Asynchronous Microphones under Existence of Source Movements with Stationary Time Periods Detection

ICASSP 2019accepted

In this paper, we propose a method of estimating the sampling frequency mismatch among asynchronous recording devices, even when the sources sometimes move. For a spatially stationary source, there is a method of estimating the sampling frequency mismatch, which appears in the drift of the time diff…

Cited by 0SourceScholar
2019

Mask-based MVDR Beamformer for Noisy Multisource Environments: Introduction of Time-varying Spatial Covariance Model

ICASSP 2019accepted

This paper proposes a method for designing a time-varying minimum variance distortionless response (MVDR) beamformer using time-frequency masks, with the aim of improving speech enhancement in noisy multi-speaker environments. A key to successful beamforming is to estimate accurately a time-varying…

Cited by 0SourceScholar
2019

Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders

ICASSP 2019accepted

We introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speec…

Cited by 0SourceScholar
2018

Listening to Each Speaker One by One with Recurrent Selective Hearing Networks

ICASSP 2018accepted

Deep learning-based single-channel source separation algorithms are currently being actively investigated. Among them, Deep Clustering (DC) and Deep Attractor Networks (DANs) have made it possible to separate an arbitrary number of speakers. In particular, they cleverly combine a neural network and…

Cited by 0SourceScholar
2018

Meeting Recognition with Asynchronous Distributed Microphone Array Using Block-Wise Refinement of Mask-Based MVDR Beamformer

ICASSP 2018accepted

This paper addresses a front-end system for speech recognition of spontaneous conversational speech signals that are recorded with asynchronous distributed microphones such as smartphones. In our previous work, we proposed combining blind synchronization and a state-of-the-art microphone array speec…

Cited by 0SourceScholar
2018

Optimization of Speaker-Aware Multichannel Speech Extraction with ASR Criterion

ICASSP 2018accepted

This paper addresses the problem of recognizing speech corrupted by overlapping speakers in a multichannel setting. To extract a target speaker from the mixture, we use a neural network based beamformer which uses masks estimated by a neural network to compute statistically optimal spatial filters.…

Cited by 0SourceScholar
2018

Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier Model

ICASSP 2018accepted

This paper proposes a new model for accurately rescoring (reranking) N-best speech recognition hypothesis lists. The model is based on state-of-the-art neural networks (NNs) and provides the minimum necessary functionality to perform N-best rescoring, i.e. one-on-one hypothesis comparison on a given…

Cited by 0SourceScholar
2018

Sequence Training of Encoder-Decoder Model Using Policy Gradient for End-to-End Speech Recognition

ICASSP 2018accepted

The standard evaluation metric of automatic speech recognition (ASR) is the word error rate (WER), which measures the dissimilarity between recognized word sequences and their ground truth. Many training algorithms designed to reduce sequence-level errors such as WER have been proposed for hidden Ma…

Cited by 0SourceScholar
2018

Single Channel Target Speaker Extraction and Recognition with Speaker Beam

ICASSP 2018accepted

This paper addresses the problem of single channel speech recognition of a target speaker in a mixture of speech signals. We propose to exploit auxiliary speaker information provided by an adaptation utterance from the target speaker to extract and recognize only that speaker. Using such auxiliary i…

Cited by 0SourceScholar
2017

Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models

ICASSP 2017accepted

Adapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount…

Cited by 0SourceScholar
2017

Deep mixture density network for statistical model-based feature enhancement

ICASSP 2017accepted

We propose a novel framework designed to extend conventional deep neural network (DNN)-based feature enhancement approaches. In general, the conventional DNN-based feature enhancement framework aims to map input noisy observation to clean speech or a binary/ soft mask in a deterministic way, assumin…

Cited by 0SourceScholar
2017

Feedback connection for deep neural network-based acoustic modeling

ICASSP 2017accepted

The use of auxiliary features is an effective way to improve the performance of deep neural network (DNN)-based acoustic models. Most approaches use auxiliary features that represent the speaker or the environment. These auxiliary features are usually computed independently of the acoustic model. Th…

Cited by 0SourceScholar
2017

Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features

ICASSP 2017accepted

We propose a new concept for adapting CNN-based acoustic models using spatial diffuseness features as auxiliary information about the acoustic environment: the spatial diffuseness features are simultaneously employed as acoustic-model input features and to estimate environmental cues for context ada…

Cited by 0SourceScholar
2017

Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environments

ICASSP 2017accepted

Here we propose online adaptive beamforming for automatic speech recognition (ASR) in meetings in noisy, reverberant environments. The proposed method is based on recently developed mask-based beamforming, in which accurate mask estimation and diarization are paramount. Real-world experiments have s…

Cited by 0SourceScholar
2016

Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditions

ICASSP 2016accepted

Deep neural network (DNN) based acoustic models have greatly improved the performance of automatic speech recognition (ASR) for various tasks. Further performance improvements have been reported when making DNNs aware of the acoustic context (e.g. speaker or environment) for example by adding auxili…

Cited by 0SourceScholar
2016

Joint acoustic factor learning for robust deep neural network based automatic speech recognition

ICASSP 2016accepted

Deep neural networks (DNNs) for acoustic modeling have been shown to provide impressive results on many state-of-the-art automatic speech recognition (ASR) applications. However, DNN performance degrades due to mismatches in training and testing conditions and thus adaptation is necessary. In this p…

Cited by 0SourceScholar
2015

Context adaptive deep neural networks for fast acoustic model adaptation

ICASSP 2015accepted

Deep neural networks (DNNs) are widely used for acoustic modeling in automatic speech recognition (ASR), since they greatly outperform legacy Gaussian mixture model-based systems. However, the levels of performance achieved by current DNN-based systems remain far too low in many tasks, e.g. when the…

Cited by 0SourceScholar
2015

Exploring multi-channel features for denoising-autoencoder-based speech enhancement

ICASSP 2015accepted

This paper investigates a multi-channel denoising autoencoder (DAE)-based speech enhancement approach. In recent years, deep neural network (DNN)-based monaural speech enhancement and robust automatic speech recognition (ASR) approaches have attracted much attention due to their high performance. Al…

Cited by 0SourceScholar
2015

WFST-based structural classification integrating dnn acoustic features and RNN language features for speech recognition

ICASSP 2015accepted

This paper proposes a method to train Weighted Finite State Transducer (WFST) based structural classifiers using deep neural network (DNN) acoustic features and recurrent neural network (RNN) language features for speech recognition. Structural classification is an effective approach to achieve high…

Cited by 0SourceScholar