← Search

Yosuke Higuchi

12 accepted papers

2025

Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model for Guiding End-to-End Speech Recognition

ICASSP 2025accepted

We propose to utilize an instruction-tuned large language model (LLM) for guiding the text generation process in automatic speech recognition (ASR). Modern LLMs are adept at performing various text generation tasks through zero-shot learning, prompted with instructions designed for specific objectiv…

Cited by 0SourceScholar
2024

Parody Detection Using Source-Target Attention with Teacher-Forced Lyrics

ICASSP 2024accepted

We propose an approach to detect parodies in singing voices, analyzing attention weights derived from an encoder-decoder-based automatic speech recognition (ASR) model. Here, parodies involve modifying and singing existing lyrics written for songs. Sharing such modified singing voices on the interne…

Cited by 0SourceScholar
2023

BECTRA: Transducer-Based End-To-End ASR with Bert-Enhanced Encoder

ICASSP 2023accepted

We present BERT-CTC-Transducer (BECTRA), a novel end-to-end automatic speech recognition (E2E-ASR) model formulated by the transducer with a BERT-enhanced encoder. Integrating a large-scale pre-trained language model (LM) into E2E-ASR has been actively studied, aiming to utilize versatile linguistic…

Cited by 0SourceScholar
2023

Intermpl: Momentum Pseudo-Labeling With Intermediate CTC Loss

ICASSP 2023accepted

This paper presents InterMPL, a semi-supervised learning method of end-to-end automatic speech recognition (ASR) that performs pseudo-labeling (PL) with intermediate supervision. Momentum PL (MPL) trains a connectionist temporal classification (CTC)-based model on unlabeled data by continuously gene…

Cited by 0SourceScholar
2022

Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy

ICASSP 2022accepted

Pseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to enhance the performance of end-to-end automatic speech recognition (ASR). Our prior work proposed momentum pseudo-labelin…

Cited by 14SourceScholar
2022

BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model

EMNLP 2022finding

This paper presents BERT-CTC, a novel formulation of end-to-end speech recognition that adapts BERT for connectionist temporal classification (CTC). Our formulation relaxes the conditional independence assumptions used in conventional CTC and incorporates linguistic knowledge through the explicit ou…

2022

Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword Units

ICASSP 2022accepted

In end-to-end automatic speech recognition (ASR), a model is expected to implicitly learn representations suitable for recognizing a word-level sequence. However, the huge abstraction gap between input acoustic signals and output linguistic tokens makes it challenging for a model to learn the repres…

Cited by 0SourceScholar
2022

Improving Non-Autoregressive End-to-End Speech Recognition with Pre-Trained Acoustic and Language Models

ICASSP 2022accepted

While Transformers have achieved promising results in end-to-end (E2E) automatic speech recognition (ASR), their autoregressive (AR) structure becomes a bottleneck for speeding up the decoding process. For real-world deployment, ASR systems are desired to be highly accurate while achieving fast infe…

Cited by 0SourceScholar
2021

Improved Mask-CTC for Non-Autoregressive End-to-End ASR

ICASSP 2021accepted

For real-world deployment of automatic speech recognition (ASR), the system is desired to be capable of fast inference while relieving the requirement of computational resources. The recently proposed end-to-end ASR system based on mask-predict with connectionist temporal classification (CTC), Mask-…

Cited by 0SourceScholar
2021

ORTHROS: non-autoregressive end-to-end speech translation With dual-decoder

ICASSP 2021accepted

Fast inference speed is an important goal towards real-world deployment of speech translation (ST) systems. End-to-end (E2E) models based on the encoder-decoder architecture are more suitable for this goal than traditional cascaded systems, but their effectiveness regarding decoding speed has not be…

Cited by 0SourceScholar
2021

Recent Developments on Espnet Toolkit Boosted By Conformer

ICASSP 2021accepted

In this study, we present recent developments on ESPnet: End-to- End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a wide range of end- to-end speech processing applications, suc…

Cited by 0SourceScholar