← Search

Tetsuji Ogawa

19 accepted papers

2025

Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model for Guiding End-to-End Speech Recognition

ICASSP 2025accepted

We propose to utilize an instruction-tuned large language model (LLM) for guiding the text generation process in automatic speech recognition (ASR). Modern LLMs are adept at performing various text generation tasks through zero-shot learning, prompted with instructions designed for specific objectiv…

Cited by 0SourceScholar
2024

Parody Detection Using Source-Target Attention with Teacher-Forced Lyrics

ICASSP 2024accepted

We propose an approach to detect parodies in singing voices, analyzing attention weights derived from an encoder-decoder-based automatic speech recognition (ASR) model. Here, parodies involve modifying and singing existing lyrics written for songs. Sharing such modified singing voices on the interne…

Cited by 0SourceScholar
2023

BECTRA: Transducer-Based End-To-End ASR with Bert-Enhanced Encoder

ICASSP 2023accepted

We present BERT-CTC-Transducer (BECTRA), a novel end-to-end automatic speech recognition (E2E-ASR) model formulated by the transducer with a BERT-enhanced encoder. Integrating a large-scale pre-trained language model (LM) into E2E-ASR has been actively studied, aiming to utilize versatile linguistic…

Cited by 0SourceScholar
2023

Conversation-Oriented ASR with Multi-Look-Ahead CBS Architecture

ICASSP 2023accepted

During conversations, humans are capable of inferring the intention of the speaker at any point of the speech to prepare the following action promptly. Such ability is also the key for conversational systems to achieve rhythmic and natural conversation. To perform this, the automatic speech recognit…

Cited by 0SourceScholar
2023

Intermpl: Momentum Pseudo-Labeling With Intermediate CTC Loss

ICASSP 2023accepted

This paper presents InterMPL, a semi-supervised learning method of end-to-end automatic speech recognition (ASR) that performs pseudo-labeling (PL) with intermediate supervision. Momentum PL (MPL) trains a connectionist temporal classification (CTC)-based model on unlabeled data by continuously gene…

Cited by 1SourceScholar
2023

Neural Diarization with Non-Autoregressive Intermediate Attractors

ICASSP 2023accepted

End-to-end neural diarization (EEND) with encoder-decoder-based attractors (EDA) is a promising method to handle the whole speaker diarization problem simultaneously with a single neural network. While the EEND model can produce all frame-level speaker labels simultaneously, it disregards output lab…

Cited by 14SourceScholar
2022

BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model

EMNLP 2022finding

This paper presents BERT-CTC, a novel formulation of end-to-end speech recognition that adapts BERT for connectionist temporal classification (CTC). Our formulation relaxes the conditional independence assumptions used in conventional CTC and incorporates linguistic knowledge through the explicit ou…

2022

Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword Units

ICASSP 2022accepted

In end-to-end automatic speech recognition (ASR), a model is expected to implicitly learn representations suitable for recognizing a word-level sequence. However, the huge abstraction gap between input acoustic signals and output linguistic tokens makes it challenging for a model to learn the repres…

Cited by 0SourceScholar
2022

Remix-Cycle-Consistent Learning on Adversarially Learned Separator for Accurate and Stable Unsupervised Speech Separation

ICASSP 2022accepted

A new learning algorithm for speech separation networks is designed to explicitly reduce residual noise and artifacts in the separated signal in an unsupervised manner. Generative adversarial networks are known to be effective in constructing separation networks when the ground truth for the observe…

Cited by 0SourceScholar
2021

Improved Mask-CTC for Non-Autoregressive End-to-End ASR

ICASSP 2021accepted

For real-world deployment of automatic speech recognition (ASR), the system is desired to be capable of fast inference while relieving the requirement of computational resources. The recently proposed end-to-end ASR system based on mask-predict with connectionist temporal classification (CTC), Mask-…

Cited by 0SourceScholar
2020

Deep Speech Extraction with Time-Varying Spatial Filtering Guided By Desired Direction Attractor

ICASSP 2020accepted

In this investigation, a deep neural network (DNN) based speech extraction method is proposed to enhance a speech signal propagating from the desired direction. The proposed method integrates knowledge based on a sound propagation model and the time-varying characteristics of a speech source, into a…

Cited by 0SourceScholar
2020

Exploiting Narrative Context and A Priori Knowledge of Categories in Textual Emotion Classification

COLING 2020main

Recognition of the mental state of a human character in text is a major challenge in natural language processing. In this study, we investigate the efficacy of the narrative context in recognizing the emotional states of human characters in text and discuss an approach to make use of a priori knowle…

Cited by 4SourcePDFScholar
2020

Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances

ICASSP 2020accepted

This paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and tes…

Cited by 0SourceScholar
2019

Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware Training

ICASSP 2019accepted

An adversarial denoising autoencoder (ADAE) with noise-aware training is proposed and successfully applied to post-filtering for linear noise reduction. The ADAE is effective for attenuating interference sounds, however, it is difficult to learn to handle its various unexpected harmful effects (e.g.…

Cited by 2SourceScholar
2018

Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations

ICASSP 2018accepted

Training recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to…

Cited by 0SourceScholar
2018

Speaker Invariant Feature Extraction for Zero-Resource Languages with Adversarial Learning

ICASSP 2018accepted

We introduce a novel type of representation learning to obtain a speaker invariant feature for zero-resource languages. Speaker adaptation is an important technique to build a robust acoustic model. For a zero-resource language, however, conventional model-dependent speaker adaptation methods such a…

Cited by 0SourceScholar
2015

A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditions

ICASSP 2015accepted

The present paper dealt with speaker clustering for speech corrupted by noise. In general, the performance of speaker clustering significantly depends on how well the similarities between speech utterances can be measured. The recently proposed i-vector-based cosine similarity has yielded the state-…

Cited by 8SourceScholar
2015

Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial Workshop

ICASSP 2015accepted

A group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which…

Cited by 0SourceScholar