← Search

Atsunori Ogawa

29 accepted papers

2025

Bridging Speech and Text Foundation Models with ReShape Attention

ICASSP 2025accepted

This paper investigates cascade approaches bridging speech and text foundation models (FMs) for speech translation (ST). We address the limitations of cascade systems which suffer from the propagation of speech recognition errors and the lack of access to acoustic information. We propose a ReShape A…

Cited by 0SourceScholar
2025

Speech Emotion Recognition Based on Large-Scale Automatic Speech Recognizer

ICASSP 2025accepted

This paper proposes a novel speech emotion recognition (SER) method that fully leverages the architecture of Whisper, a large-scale automatic speech recognition (ASR) model. The conventional SER models using a pre-trained speech encoder may fail to capture linguistic content since their decoders are…

Cited by 0SourceScholar
2024

NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization

ICASSP 2024accepted

This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)based dereverberation as a front end, and separately applies end-to-end neural diarization with vector clustering…

Cited by 0SourceScholar
2024

Train Long and Test Long: Leveraging Full Document Contexts in Speech Processing

ICASSP 2024accepted

The quadratic memory complexity of self-attention has generally restricted Transformer-based models to utterance-based speech processing, preventing models from leveraging long-form contexts. A common solution has been to formulate long-form speech processing into a streaming problem, only using lim…

Cited by 0SourceScholar
2023

Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition

ICASSP 2023accepted

We propose a new shallow fusion (SF) method to exploit an external backward language model (BLM) for end-to-end automatic speech recognition (ASR). The BLM has complementary characteristics with a forward language model (FLM), and the effectiveness of their combination has been confirmed by rescorin…

Cited by 0SourceScholar
2023

Leveraging Large Text Corpora For End-To-End Speech Summarization

ICASSP 2023accepted

End-to-end speech summarization (E2E SSum) is a technique to directly generate summary sentences from speech. Compared with the cascade approach, which combines automatic speech recognition (ASR) and text summarization models, the E2E approach is more promising because it mitigates ASR errors, incor…

Cited by 0SourceScholar
2023

Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text Encoders

ICASSP 2023accepted

Speech summarization requires processing several minute-long speech sequences to allow exploiting the whole context of a spoken document. A conventional approach is a cascade of automatic speech recognition (ASR) and text summarization (TS). However, the cascade systems are sensitive to ASR errors.…

Cited by 0SourceScholar
2022

Integrating Multiple ASR Systems into NLP Backend with Attention Fusion

ICASSP 2022accepted

Spoken language processing (SLP) systems such as speech summarization and translation can be achieved by cascade models. It combines an automatic speech recognition (ASR) frontend and a natural language processing (NLP) backend including machine translation (MT) or text summarization (TS). With this…

Cited by 0SourceScholar
2022

Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models

ICASSP 2022accepted

We investigate the effectiveness of using a large ensemble of advanced neural language models (NLMs) for lattice rescoring on automatic speech recognition (ASR) hypotheses. Previous studies have reported the effectiveness of combining a small number of NLMs. In contrast, in this study, we combine up…

Cited by 0SourceScholar
2021

Age-VOX-Celeb: Multi-Modal Corpus for Facial and Speech Estimation

ICASSP 2021accepted

Estimating a speaker’s age from their speech is more challenging than age estimation from their face because of insufficiently available public corpora. To tackle this problem, we construct a new audio-visual age corpus named AgeVoxCeleb by annotating age labels to VoxCeleb2 videos. AgeVoxCeleb is t…

Cited by 0SourceScholar
2021

BLSTM-Based Confidence Estimation for End-to-End Speech Recognition

ICASSP 2021accepted

Confidence estimation, in which we estimate the reliability of each recognized token (e.g., word, sub-word, and character) in automatic speech recognition (ASR) hypotheses and detect incorrectly recognized tokens, is an important function for developing ASR applications. In this study, we perform co…

Cited by 0SourceScholar
2020

Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances

ICASSP 2020accepted

This paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and tes…

Cited by 0SourceScholar
2020

Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster Information

ICASSP 2020accepted

This paper proposes a general post-processing method for improving speaker-attribute estimation. Estimating speaker-specific attributes such as age and gender is an important task with a wide range of applications. While the recent proposed deep neural network-based end-to-end approach achieves high…

Cited by 0SourceScholar
2019

A Unified Framework for Feature-based Domain Adaptation of Neural Network Language Models

ICASSP 2019accepted

An important task for language models is the adaptation of general-domain models to specific target domains. For neural network-based language models, feature-based domain adaptation has been a popular method in previous research. Conventional methods use an adaptation feature providing context info…

Cited by 0SourceScholar
2019

A Unified Framework for Neural Speech Separation and Extraction

ICASSP 2019accepted

The development of deep learning techniques has triggered the active investigation of neural network-based speech enhancement approaches. In particular, single-channel blind (uninformed) speech separation and speaker-aware (informed) speech extraction have received increased interest. Blind speech s…

Cited by 21SourceScholar
2019

ILP-based Compressive Speech Summarization with Content Word Coverage Maximization and Its Oracle Performance Analysis

ICASSP 2019accepted

We propose an integer linear programming (ILP)-based compressive speech summarization method that maximizes the coverage of content words in a resultant summary. It is an unsupervised method and, under the designed constraints, it performs a single-step globally optimal summarization of a given long…

Cited by 0SourceScholar
2019

Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders

ICASSP 2019accepted

We introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speec…

Cited by 44SourceScholar
2018

Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations

ICASSP 2018accepted

Training recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to…

Cited by 0SourceScholar
2018

Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier Model

ICASSP 2018accepted

This paper proposes a new model for accurately rescoring (reranking) N-best speech recognition hypothesis lists. The model is based on state-of-the-art neural networks (NNs) and provides the minimum necessary functionality to perform N-best rescoring, i.e. one-on-one hypothesis comparison on a given…

Cited by 0SourceScholar
2018

Sequence Training of Encoder-Decoder Model Using Policy Gradient for End-to-End Speech Recognition

ICASSP 2018accepted

The standard evaluation metric of automatic speech recognition (ASR) is the word error rate (WER), which measures the dissimilarity between recognized word sequences and their ground truth. Many training algorithms designed to reduce sequence-level errors such as WER have been proposed for hidden Ma…

Cited by 0SourceScholar
2018

Single Channel Target Speaker Extraction and Recognition with Speaker Beam

ICASSP 2018accepted

This paper addresses the problem of single channel speech recognition of a target speaker in a mixture of speech signals. We propose to exploit auxiliary speaker information provided by an adaptation utterance from the target speaker to extract and recognize only that speaker. Using such auxiliary i…

Cited by 0SourceScholar
2017

Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models

ICASSP 2017accepted

Adapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount…

Cited by 0SourceScholar
2017

Deep mixture density network for statistical model-based feature enhancement

ICASSP 2017accepted

We propose a novel framework designed to extend conventional deep neural network (DNN)-based feature enhancement approaches. In general, the conventional DNN-based feature enhancement framework aims to map input noisy observation to clean speech or a binary/ soft mask in a deterministic way, assumin…

Cited by 0SourceScholar
2017

Feedback connection for deep neural network-based acoustic modeling

ICASSP 2017accepted

The use of auxiliary features is an effective way to improve the performance of deep neural network (DNN)-based acoustic models. Most approaches use auxiliary features that represent the speaker or the environment. These auxiliary features are usually computed independently of the acoustic model. Th…

Cited by 0SourceScholar
2017

Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features

ICASSP 2017accepted

We propose a new concept for adapting CNN-based acoustic models using spatial diffuseness features as auxiliary information about the acoustic environment: the spatial diffuseness features are simultaneously employed as acoustic-model input features and to estimate environmental cues for context ada…

Cited by 0SourceScholar
2016

Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditions

ICASSP 2016accepted

Deep neural network (DNN) based acoustic models have greatly improved the performance of automatic speech recognition (ASR) for various tasks. Further performance improvements have been reported when making DNNs aware of the acoustic context (e.g. speaker or environment) for example by adding auxili…

Cited by 36SourceScholar
2016

Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognition

ICASSP 2016accepted

This paper addresses a minimum variance distortionless response (MVDR) beamforming based speech enhancement approach for meeting speech recognition. In a meeting situation, speaker overlaps and noise signals are not negligible. To handle these issues, we employ MVDR beamforming, where accurate estim…

Cited by 0SourceScholar
2015

ASR error detection and recognition rate estimation using deep bidirectional recurrent neural networks

ICASSP 2015accepted

Recurrent neural networks (RNNs) have recently been applied as the classifiers for sequential labeling problems. In this paper, deep bidirectional RNNs (DBRNNs) are applied for the first time to error detection in automatic speech recognition (ASR), which is a sequential labeling problem. We investi…

Cited by 0SourceScholar
2015

Double-layer neighborhood graph based similarity search for fast query-by-example spoken term detection

ICASSP 2015accepted

This paper presents a novel double-layer neighborhood graph index for acceleration of similarity search that accomplishes fast querybyexample spoken term detection (STD). When a query segment is given, our proposed STD method finds similar segments to the query from an utterance data set by efficien…

Cited by 0SourceScholar