← Search

Tetsunori Kobayashi

16 accepted papers

2025

Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model for Guiding End-to-End Speech Recognition

ICASSP 2025accepted

We propose to utilize an instruction-tuned large language model (LLM) for guiding the text generation process in automatic speech recognition (ASR). Modern LLMs are adept at performing various text generation tasks through zero-shot learning, prompted with instructions designed for specific objectiv…

Cited by 0SourceScholar
2023

BECTRA: Transducer-Based End-To-End ASR with Bert-Enhanced Encoder

ICASSP 2023accepted

We present BERT-CTC-Transducer (BECTRA), a novel end-to-end automatic speech recognition (E2E-ASR) model formulated by the transducer with a BERT-enhanced encoder. Integrating a large-scale pre-trained language model (LM) into E2E-ASR has been actively studied, aiming to utilize versatile linguistic…

Cited by 0SourceScholar
2023

Conversation-Oriented ASR with Multi-Look-Ahead CBS Architecture

ICASSP 2023accepted

During conversations, humans are capable of inferring the intention of the speaker at any point of the speech to prepare the following action promptly. Such ability is also the key for conversational systems to achieve rhythmic and natural conversation. To perform this, the automatic speech recognit…

Cited by 0SourceScholar
2023

Intermpl: Momentum Pseudo-Labeling With Intermediate CTC Loss

ICASSP 2023accepted

This paper presents InterMPL, a semi-supervised learning method of end-to-end automatic speech recognition (ASR) that performs pseudo-labeling (PL) with intermediate supervision. Momentum PL (MPL) trains a connectionist temporal classification (CTC)-based model on unlabeled data by continuously gene…

Cited by 0SourceScholar
2022

BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model

EMNLP 2022finding

This paper presents BERT-CTC, a novel formulation of end-to-end speech recognition that adapts BERT for connectionist temporal classification (CTC). Our formulation relaxes the conditional independence assumptions used in conventional CTC and incorporates linguistic knowledge through the explicit ou…

2022

Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword Units

ICASSP 2022accepted

In end-to-end automatic speech recognition (ASR), a model is expected to implicitly learn representations suitable for recognizing a word-level sequence. However, the huge abstraction gap between input acoustic signals and output linguistic tokens makes it challenging for a model to learn the repres…

Cited by 0SourceScholar
2022

Phrase-Level Localization of Inconsistency Errors in Summarization by Weak Supervision

COLING 2022main

Although the fluency of automatically generated abstractive summaries has improved significantly with advanced methods, the inconsistency that remains in summarization is recognized as an issue to be addressed. In this study, we propose a methodology for localizing inconsistency errors in summarizat…

2021

Improved Mask-CTC for Non-Autoregressive End-to-End ASR

ICASSP 2021accepted

For real-world deployment of automatic speech recognition (ASR), the system is desired to be capable of fast inference while relieving the requirement of computational resources. The recently proposed end-to-end ASR system based on mask-predict with connectionist temporal classification (CTC), Mask-…

Cited by 0SourceScholar
2020

Deep Speech Extraction with Time-Varying Spatial Filtering Guided By Desired Direction Attractor

ICASSP 2020accepted

In this investigation, a deep neural network (DNN) based speech extraction method is proposed to enhance a speech signal propagating from the desired direction. The proposed method integrates knowledge based on a sound propagation model and the time-varying characteristics of a speech source, into a…

Cited by 0SourceScholar
2020

Exploiting Narrative Context and A Priori Knowledge of Categories in Textual Emotion Classification

COLING 2020main

Recognition of the mental state of a human character in text is a major challenge in natural language processing. In this study, we investigate the efficacy of the narrative context in recognizing the emotional states of human characters in text and discuss an approach to make use of a priori knowle…

Cited by 4SourcePDFScholar
2020

Sentiment Analysis for Emotional Speech Synthesis in a News Dialogue System

COLING 2020main

As smart speakers and conversational robots become ubiquitous, the demand for expressive speech synthesis has increased. In this paper, to control the emotional parameters of the speech synthesis according to certain dialogue contents, we construct a news dataset with emotion labels (“positive,” “ne…

2019

Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware Training

ICASSP 2019accepted

An adversarial denoising autoencoder (ADAE) with noise-aware training is proposed and successfully applied to post-filtering for linear noise reduction. The ADAE is effective for attenuating interference sounds, however, it is difficult to learn to handle its various unexpected harmful effects (e.g.…

Cited by 0SourceScholar
2018

Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations

ICASSP 2018accepted

Training recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to…

Cited by 0SourceScholar
2018

Speaker Invariant Feature Extraction for Zero-Resource Languages with Adversarial Learning

ICASSP 2018accepted

We introduce a novel type of representation learning to obtain a speaker invariant feature for zero-resource languages. Speaker adaptation is an important technique to build a robust acoustic model. For a zero-resource language, however, conventional model-dependent speaker adaptation methods such a…

Cited by 0SourceScholar
2015

A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditions

ICASSP 2015accepted

The present paper dealt with speaker clustering for speech corrupted by noise. In general, the performance of speaker clustering significantly depends on how well the similarities between speech utterances can be measured. The recently proposed i-vector-based cosine similarity has yielded the state-…

Cited by 0SourceScholar