← Search

Kohei Matsuura

10 accepted papers

2025

Advancing Streaming ASR with Chunk-wise Attention and Trans-chunk Selective State Spaces

ICASSP 2025accepted

This paper explores enhancing streaming speech recognition through the integration of chunk-wise attention and selective state space models (SSMs). The proposed framework replaces the quadratic complexity of attention-based context incorporation with a fully recurrent module based on selective SSMs.…

Cited by 0SourceScholar
2025

Alignment-Free Training for Transducer-based Multi-Talker ASR

ICASSP 2025accepted

Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to achieve recognition without relying on costly front-end source separation. MT-RNNT is conventionally implemented using arch…

Cited by 0SourceScholar
2025

Bridging Speech and Text Foundation Models with ReShape Attention

ICASSP 2025accepted

This paper investigates cascade approaches bridging speech and text foundation models (FMs) for speech translation (ST). We address the limitations of cascade systems which suffer from the propagation of speech recognition errors and the lack of access to acoustic information. We propose a ReShape A…

Cited by 0SourceScholar
2024

What Do Self-Supervised Speech and Speaker Models Learn? New Findings from a Cross Model Layer-Wise Analysis

ICASSP 2024accepted

Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representations. In contrast, speaker SSL models, exemplified by DINO-based models, adopt utt…

Cited by 0SourceScholar
2023

Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models

ICASSP 2023accepted

Self-supervised learning (SSL) has been dramatically successful not only in monolingual but also in cross-lingual settings. However, since the two settings have been studied individually in general, there has been little research focusing on how effective a cross-lingual model is in comparison with…

Cited by 5SourceScholar
2023

Improving Scheduled Sampling for Neural Transducer-Based ASR

ICASSP 2023accepted

The recurrent neural network-transducer (RNNT) is a promising approach for automatic speech recognition (ASR) with the introduction of a prediction network that autoregressively considers linguistic aspects. To train the autoregressive part, the ground-truth tokens are used as substitutions for the…

Cited by 0SourceScholar
2023

Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning

ICASSP 2023accepted

In this paper, we propose novel cross-lingual self-supervised speech representation learning methods that explicitly consider language information. Cross-lingual self-supervised speech representation learning has been studied to make effective use of diverse data in various languages. Previous metho…

Cited by 0SourceScholar
2023

Leveraging Large Text Corpora For End-To-End Speech Summarization

ICASSP 2023accepted

End-to-end speech summarization (E2E SSum) is a technique to directly generate summary sentences from speech. Compared with the cascade approach, which combines automatic speech recognition (ASR) and text summarization models, the E2E approach is more promising because it mitigates ASR errors, incor…

Cited by 0SourceScholar
2023

Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text Encoders

ICASSP 2023accepted

Speech summarization requires processing several minute-long speech sequences to allow exploiting the whole context of a spoken document. A conventional approach is a cascade of automatic speech recognition (ASR) and text summarization (TS). However, the cascade systems are sensitive to ASR errors.…

Cited by 12SourceScholar
2022

Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration

ICASSP 2022accepted

In this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (…

Cited by 0SourceScholar