← Search

Takashi Fukuda

6 accepted papers

2025

Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASR

ICASSP 2025accepted

There is an on-going body of research on training separate dedicated models for either short-form or long-form utterances. Multi-form acoustic models that are simply trained on combined data from long-form and short-form utterances often suffer from various negative impacts due to the diversity of a…

Cited by 0SourceScholar
2023

Effective Training of RNN Transducer Models on Diverse Sources of Speech and Text Data

ICASSP 2023accepted

This paper proposes a novel modeling framework for effective training of end-to-end automatic speech recognition (ASR) models on various sources of data from diverse domains: speech paired with clean ground truth transcripts, speech with noisy pseudo transcripts from semi-supervised decodes and unpa…

Cited by 0SourceScholar
2021

Generalized Knowledge Distillation from an Ensemble of Specialized Teachers Leveraging Unsupervised Neural Clustering

ICASSP 2021accepted

This paper proposes an improved generalized knowledge distillation framework with multiple dissimilar teacher networks, each of which is specialized for a specific domain, to make a deployable student network more robust to challenging acoustic environments. In this paper, we first address a method…

Cited by 0SourceScholar
2017

Effective joint training of denoising feature space transforms and Neural Network based acoustic models

ICASSP 2017accepted

Neural Network (NN) based acoustic frontends, such as denoising autoencoders, are actively being investigated to improve the robustness of NN based acoustic models to various noise conditions. In recent work the joint training of such frontends with backend NNs has been shown to significantly improv…

Cited by 0SourceScholar
2017

Harmonic feature fusion for robust neural network-based acoustic modeling

ICASSP 2017accepted

Acoustic modeling with deep learning has drastically improved the performance of automatic speech recognition (ASR) where the main stream of the acoustic feature is still log-Mel filtered one. While the log-Mel filtered features lose harmonic-structure information, they still include useful informat…

Cited by 0SourceScholar
2016

Convolutional neural network pre-trained with projection matrices on linear discriminant analysis

ICASSP 2016accepted

Recently, the hybrid architecture of a neural network (NN) and a hidden Markov model (HMM) has shown significant improvement on automatic speech recognition (ASR) over the conventional Gaussian mixture model (GMM)-based system. The convolutional neural network (CNN), a successful NN-based system, ca…

Cited by 0SourceScholar