← Search

Michiel Bacchiani

9 accepted papers

2022

Knowledge Transfer from Large-Scale Pretrained Language Models to End-To-End Speech Recognizers

ICASSP 2022accepted

End-to-end speech recognition is a promising technology for enabling compact automatic speech recognition (ASR) systems since it can unify the acoustic and language model into a single neural network. However, as a drawback, training of end-to-end speech recognizers always requires transcribed utter…

Cited by 0SourceScholar
2018

Multi-Dialect Speech Recognition with a Single Sequence-to-Sequence Model

ICASSP 2018accepted

Sequence-to-sequence models provide a simple and elegant solution for building speech recognition systems by folding separate components of a typical system, namely acoustic (AM), pronunciation (PM) and language (LM) models into a single neural network. In this work, we look at one such sequence-to-…

Cited by 0SourceScholar
2018

Performance of Mask Based Statistical Beamforming in a Smart Home Scenario

ICASSP 2018accepted

Mask based statistical beamforming, where signal statistics for the target and the interference gained from masking are used for beamforming, has shown great effectiveness in the two recent CHiME challenges. This idea has sparked interest in the research community and resulted in numerous proposed a…

Cited by 0SourceScholar
2018

Sampled Connectionist Temporal Classification

ICASSP 2018accepted

This article introduces and evaluates Sampled Connectionist Temporal Classification (CTC) which connects the CTC criterion to the Cross Entropy (CE) objective through sampling. Instead of computing the logarithm of the sum of the alignment path likelihoods, at each training step the sampled CTC only…

Cited by 0SourceScholar
2018

Sound Source Separation Using Phase Difference and Reliable Mask Selection Selection

ICASSP 2018accepted

In this paper, we present an algorithm called Reliable Mask Selection-Phase Difference Channel Weighting (RMS-PDCW) which selects the target source masked by a noise source using the Angle of Arrival (AoA) information calculated using the phase difference information. The RMS-PDCW algorithm selects…

Cited by 0SourceScholar
2018

Spectral Distortion Model for Training Phase-Sensitive Deep-Neural Networks for Far-Field Speech Recognition

ICASSP 2018accepted

In this paper, we present an algorithm which introduces phase-perturbation to the training database when training phase-sensitive deep neural-network models. Traditional features such as log-mel or cepstral features do not have have any phase-relevant information. However features such as raw-wavefo…

Cited by 3SourceScholar
2018

State-of-the-Art Speech Recognition with Sequence-to-Sequence Models

ICASSP 2018accepted

Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural network. In previous work, we have shown that such architectures ar…

Cited by 0SourceScholar
2016

Factored spatial and spectral multichannel raw waveform CLDNNs

ICASSP 2016accepted

Multichannel ASR systems commonly separate speech enhancement, including localization, beamforming and postfiltering, from acoustic modeling. Recently, we explored doing multichannel enhancement jointly with acoustic modeling, where beamforming and frequency decomposition was folded into one layer o…

Cited by 0SourceScholar