← Search

Takatomo Kano

6 accepted papers

2025

Bridging Speech and Text Foundation Models with ReShape Attention

ICASSP 2025accepted

This paper investigates cascade approaches bridging speech and text foundation models (FMs) for speech translation (ST). We address the limitations of cascade systems which suffer from the propagation of speech recognition errors and the lack of access to acoustic information. We propose a ReShape A…

Cited by 0SourceScholar
2025

Speech Emotion Recognition Based on Large-Scale Automatic Speech Recognizer

ICASSP 2025accepted

This paper proposes a novel speech emotion recognition (SER) method that fully leverages the architecture of Whisper, a large-scale automatic speech recognition (ASR) model. The conventional SER models using a pre-trained speech encoder may fail to capture linguistic content since their decoders are…

Cited by 0SourceScholar
2024

Train Long and Test Long: Leveraging Full Document Contexts in Speech Processing

ICASSP 2024accepted

The quadratic memory complexity of self-attention has generally restricted Transformer-based models to utterance-based speech processing, preventing models from leveraging long-form contexts. A common solution has been to formulate long-form speech processing into a streaming problem, only using lim…

Cited by 0SourceScholar
2023

Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text Encoders

ICASSP 2023accepted

Speech summarization requires processing several minute-long speech sequences to allow exploiting the whole context of a spoken document. A conventional approach is a cascade of automatic speech recognition (ASR) and text summarization (TS). However, the cascade systems are sensitive to ASR errors.…

Cited by 12SourceScholar
2022

Integrating Multiple ASR Systems into NLP Backend with Attention Fusion

ICASSP 2022accepted

Spoken language processing (SLP) systems such as speech summarization and translation can be achieved by cascade models. It combines an automatic speech recognition (ASR) frontend and a natural language processing (NLP) backend including machine translation (MT) or text summarization (TS). With this…

Cited by 0SourceScholar
2021

BLSTM-Based Confidence Estimation for End-to-End Speech Recognition

ICASSP 2021accepted

Confidence estimation, in which we estimate the reliability of each recognized token (e.g., word, sub-word, and character) in automatic speech recognition (ASR) hypotheses and detect incorrectly recognized tokens, is an important function for developing ASR applications. In this study, we perform co…

Cited by 0SourceScholar