← Search

Gakuto Kurata

14 accepted papers

2025

Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASR

ICASSP 2025accepted

There is an on-going body of research on training separate dedicated models for either short-form or long-form utterances. Multi-form acoustic models that are simply trained on combined data from long-form and short-form utterances often suffer from various negative impacts due to the diversity of a…

Cited by 0SourceScholar
2025

LLM based Text Generation for Improved Low-resource Speech Recognition Models

ICASSP 2025accepted

Limited transcribed spoken style data is a critical bottleneck in building automatic speech recognition (ASR) systems for low-resource languages. Prompting a large language model (LLM) to paraphrase input text can generate novel text data that is constrained to be semantically similar to the source…

Cited by 0SourceScholar
2024

Multiple Representation Transfer from Large Language Models to End-to-End ASR Systems

ICASSP 2024accepted

Transferring the knowledge of large language models (LLMs) is a promising technique to incorporate linguistic knowledge into end-to-end automatic speech recognition (ASR) systems. However, existing works only transfer a single representation of LLM (e.g. the last layer of pretrained BERT), while the…

Cited by 0SourceScholar
2024

Robust ASR Error Correction with Conservative Data Filtering

EMNLP 2024industry

Error correction (EC) based on large language models is an emerging technology to enhance the performance of automatic speech recognition (ASR) systems.Generally, training data for EC are collected by automatically pairing a large set of ASR hypotheses (as sources) and their gold references (as targ…

2023

Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word Dictionaries

EMNLP 2023long main

Despite the impressive performance of ASR models on mainstream benchmarks, their performance on rare words is unsatisfactory. In enterprise settings, often a focused list of entities (such as locations, names, etc) are available which can be used to adapt the model to the terminology of specific dom…

Cited by 0SourceScholar
2021

Generalized Knowledge Distillation from an Ensemble of Specialized Teachers Leveraging Unsupervised Neural Clustering

ICASSP 2021accepted

This paper proposes an improved generalized knowledge distillation framework with multiple dissimilar teacher networks, each of which is specialized for a specific domain, to make a deployable student network more robust to challenging acoustic environments. In this paper, we first address a method…

Cited by 0SourceScholar
2021

RNN Transducer Models for Spoken Language Understanding

ICASSP 2021accepted

We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available ann…

Cited by 0SourceScholar
2020

Converting Written Language to Spoken Language with Neural Machine Translation for Language Modeling

ICASSP 2020accepted

When building a language model (LM) for spontaneous speech, the ideal situation is to have a large amount of spoken, in-domain training data. Having such abundant data, however, is not realistic. We address this problem by generating texts in spoken language from those in written language by using a…

Cited by 0SourceScholar
2019

English Broadcast News Speech Recognition by Humans and Machines

ICASSP 2019accepted

With recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broa…

Cited by 0SourceScholar
2019

Improvements to N-gram Language Model Using Text Generated from Neural Language Model

ICASSP 2019accepted

Although neural language models have emerged, n-gram language models are still used for many speech recognition tasks. This paper proposes four methods to improve n-gram language models using text generated from a recurrent neural network language model (RNNLM). First, we use multiple RNNLMs from di…

Cited by 0SourceScholar
2017

Effective joint training of denoising feature space transforms and Neural Network based acoustic models

ICASSP 2017accepted

Neural Network (NN) based acoustic frontends, such as denoising autoencoders, are actively being investigated to improve the robustness of NN based acoustic models to various noise conditions. In recent work the joint training of such frontends with backend NNs has been shown to significantly improv…

Cited by 0SourceScholar
2017

Harmonic feature fusion for robust neural network-based acoustic modeling

ICASSP 2017accepted

Acoustic modeling with deep learning has drastically improved the performance of automatic speech recognition (ASR) where the main stream of the acoustic feature is still log-Mel filtered one. While the log-Mel filtered features lose harmonic-structure information, they still include useful informat…

Cited by 0SourceScholar
2016

Speech recognition robust against speech overlapping in monaural recordings of telephone conversations

ICASSP 2016accepted

Monaural (single-channel) recording is sometimes used for telephone conversations in call centers. Generally speaking, the accuracy of automatic speech recognition of a monaural recording is worse than that of the multi-channel recording of the same conversation where each speaker's voice is separat…

Cited by 0SourceScholar