← Search

Tatsuya Kawahara

31 accepted papers

2025

A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment

IROS 2025

Turn-taking is a crucial aspect of human-robot interaction, directly influencing conversational fluidity and user engagement. While previous research has explored turn-taking models in controlled environments, their robustness in real-world settings remains underexplored. In this study, we propose a

Cited by 3SourceScholar
2025

Extending Whisper for Emotion Prediction Using Word-level Pseudo Labels

ICASSP 2025accepted

This paper extends Whisper’s automatic speech recognition (ASR) capabilities to perform speech-based emotion recognition (SER) by incorporating word-level emotion classification alongside ASR output. We generate four emotion pseudo-labels (neutral, happy, sad, angry) for each word using a pretrained…

Cited by 0SourceScholar
2025

Human-Like Embodied AI Interviewer: Employing Android ERICA in Real International Conference

COLING 2025system demonstrations

This paper introduces the human-like embodied AI interviewer which integrates android robots equipped with advanced conversational capabilities, including attentive listening, conversational repairs, and user fluency adaptation. Moreover, it can analyze and present results post-interview. We conduct…

2025

Leveraging IPA and Articulatory Features as Effective Inductive Biases for Multilingual ASR Training

ICASSP 2025accepted

In recent advancements in end-to-end ASR, large-scale self-supervised or weakly supervised models have achieved a significant milestone. However, it remains challenging to train consistently high-performing multilingual models, transferable to languages without much resource. In this study, we propo…

Cited by 0SourceScholar
2025

Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection

NAACL 2025long

In human conversations, short backchannel utterances such as “yeah” and “oh” play a crucial role in facilitating smooth and engaging dialogue.These backchannels signal attentiveness and understanding without interrupting the speaker, making their accurate prediction essential for creating more natur…

Cited by 2SourcePDFScholar
2024

Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

ICASSP 2024accepted

Diffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusion process with enhanced features estimated by a predictive SE system. However, the pipeline structure currently does not…

Cited by 0SourceScholar
2024

Enhancing Two-Stage Finetuning for Speech Emotion Recognition Using Adapters

ICASSP 2024accepted

This study investigates the effective finetuning of a pretrained model using adapters for speech emotion recognition (SER). Since emotion is related with linguistic and prosodic information and also other attributes such as gender and speaking style, a framework of multi-task learning (MTL) has been…

Cited by 0SourceScholar
2024

MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction

ICASSP 2024accepted

IEEE Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD) as we expect that MOS can be used to assess how close synthesized speech is to the natural human v…

Cited by 0SourceScholar
2024

Multilingual Turn-taking Prediction Using Voice Activity Projection

COLING 2024main

This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, le…

Cited by 9SourcePDFScholar
2024

Zero- and Few-Shot Sound Event Localization and Detection

ICASSP 2024accepted

Sound event localization and detection (SELD) systems estimate direction-of-arrival (DOA) and temporal activation for sets of target classes. Neural network (NN)-based SELD systems have performed well in various sets of target classes, but they only output the DOA and temporal activation of preset c…

Cited by 0SourceScholar
2023

Domain and Language Adaptation Using Heterogeneous Datasets for Wav2vec2.0-Based Speech Recognition of Low-Resource Language

ICASSP 2023accepted

We address the effective finetuning of a large-scale pretrained model for automatic speech recognition (ASR) of lowresource languages with only a one-hour matched dataset. The finetuning is composed of domain adaptation and language adaptation, and they are conducted by using heterogeneous datasets,…

Cited by 0SourceScholar
2023

Time-Domain Speech Enhancement Assisted by Multi-Resolution Frequency Encoder and Decoder

ICASSP 2023accepted

Time-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS [1] introduces multi-resolution STFT loss to enhance performance. However, some resolutions used for STFT contain non-stationary signals, and it is challenging to learn multi-resolution frequen…

Cited by 0SourceScholar
2022

Phone-Informed Refinement of Synthesized Mel Spectrogram for Data Augmentation in Speech Recognition

ICASSP 2022accepted

While recent end-to-end automatic speech recognition (ASR) models achieve high performance, we need to prepare an abundant amount of training data, which is a barrier to apply them to a specific domain. To mitigate the lack of training data, text-to-speech (TTS) systems have been utilized to leverag…

Cited by 0SourceScholar
2022

Selective Multi-Task Learning For Speech Emotion Recognition Using Corpora Of Different Styles

ICASSP 2022accepted

While speech emotion recognition (SER) has been actively studied, the amount and variations of training data are limited compared with speech recognition and speaker recognition tasks. Therefore, it is promising to combine multiple corpora to train a generalized SER model. However, the manner of emo…

Cited by 0SourceScholar
2021

ORTHROS: non-autoregressive end-to-end speech translation With dual-decoder

ICASSP 2021accepted

Fast inference speed is an important goal towards real-world deployment of speech translation (ST) systems. End-to-end (E2E) models based on the encoder-decoder architecture are more suitable for this goal than traditional cascaded systems, but their effectiveness regarding decoding speed has not be…

Cited by 0SourceScholar
2021

Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation

NAACL 2021long

A conventional approach to improving the performance of end-to-end speech translation (E2E-ST) models is to leverage the source transcription via pre-training and joint training with automatic speech recognition (ASR) and neural machine translation (NMT) tasks. However, since the input modalities ar…

2020

Topic-relevant Response Generation using Optimal Transport for an Open-domain Dialog System

COLING 2020main

Conventional neural generative models tend to generate safe and generic responses which have little connection with previous utterances semantically and would disengage users in a dialog system. To generate relevant responses, we propose a method that employs two types of constraints - topical const…

Cited by 7SourcePDFScholar
2019

Multi-speaker Sequence-to-sequence Speech Synthesis for Data Augmentation in Acoustic-to-word Speech Recognition

ICASSP 2019accepted

The acoustic-to-word (A2W) automatic speech recognition (ASR) realizes very fast decoding with a simple architecture and achieves state-of-the-art performance. However, the A2W model suffers from the out-of-vocabulary (OOV) word problem and cannot use text-only data to improve the language modeling…

Cited by 0SourceScholar
2019

Transfer Learning of Language-independent End-to-end ASR with Language Model Fusion

ICASSP 2019accepted

This work explores better adaptation methods to low-resource languages using an external language model (LM) under the framework of transfer learning. We first build a language-independent ASR system in a unified sequence-to-sequence (S2S) architecture with a shared vocabulary among all languages. D…

Cited by 0SourceScholar
2018

Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based Model

ICASSP 2018accepted

This paper addresses end-to-end speech recognition which directly maps acoustic features to a word sequence. The acoustic-to-word model is attractive since it does not require an external language model and an elaborate decoder, resulting in extremely simple and fast decoding. The apparent drawback…

Cited by 0SourceScholar
2018

An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition

ICASSP 2018accepted

Social signals such as laughter and fillers are often observed in natural conversation, and they play various roles in human-to-human communication. Detecting these events is useful for transcription systems to generate rich transcription and for dialogue systems to behave as we do such as synchroni…

Cited by 0SourceScholar
2018

Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid Robot

ICASSP 2018accepted

This paper addresses audio-visual signal processing for conversation analysis, which involves multi-modal behavior detection and mental-state recognition. We have investigated prediction of turn-taking by the audience in a poster session from their multi-modal behaviors, and found out that the eye-g…

Cited by 0SourceScholar
2018

Efficient Learning of Articulatory Models Based on Multi-Label Training and Label Correction for Pronunciation Learning

ICASSP 2018accepted

Articulatory feedback is effective for computer-assisted pronunciation training (CAPT) systems. This paper investigates efficient model learning methods for providing articulatory information to language learners. We first propose an articulatory attribute modeling method based on a multi-label lear…

Cited by 0SourceScholar
2018

Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization

ICASSP 2018accepted

This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Al…

Cited by 0SourceScholar
2018

Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech Recognition

ICASSP 2018accepted

This paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to…

Cited by 0SourceScholar
2017

Bayesian multichannel nonnegative matrix factorization for audio source separation and localization

ICASSP 2017accepted

This paper presents a Bayesian extension of multichannel nonnegative matrix factorization (MNMF) that decomposes the complex spectrograms of mixture signals recorded by a microphone array into basis spectra, their temporal activations, and the spatial correlation matrices of sources (directions) in…

Cited by 0SourceScholar
2017

Effective articulatory modeling for pronunciation error detection of L2 learner without non-native training data

ICASSP 2017accepted

For effective articulatory feedback in computer-assisted pronunciation training (CAPT) systems, we address effective articulatory models of second language (L2) learners' speech without using such data, which is difficult to collect and annotate in a large scale. Context-dependent articulatory attri…

Cited by 0SourceScholar
2017

Semi-supervised ensemble DNN acoustic model training

ICASSP 2017accepted

It is very important to exploit abundant unlabeled speech for improving the acoustic model training in automatic speech recognition (ASR). Semi-supervised training methods incorporate unlabeled data in addition to labeled data to enhance the model training, but it encounters the error-prone label pr…

Cited by 0SourceScholar
2016

Data selection from multiple ASR systems' hypotheses for unsupervised acoustic model training

ICASSP 2016accepted

This paper addresses unsupervised training of DNN acoustic model, by exploiting a large amount of unlabeled data with CRF-based classifiers. In the proposed scheme, we obtain ASR hypotheses by complementary GMM and DNN based ASR systems. Then, a set of dedicated classifiers are designed and trained…

Cited by 0SourceScholar
2015

Deep autoencoders augmented with phone-class feature for reverberant speech recognition

ICASSP 2015accepted

This paper addresses reverberant speech recognition based on front-end processing using DAE (Deep AutoEncoder) coupled with DNN (Deep Neural Network) acoustic model. DAE can effectively and flexibly learn mapping from corrupted speech to the original clean speech based on the deep learning scheme. W…

Cited by 0SourceScholar
2015

Language model adaptation for academic lectures using character recognition result of presentation slides

ICASSP 2015accepted

For automatic speech recognition (ASR) of lectures, texts of presentation slides are expected to be useful for adapting a language model, while slide texts are not always available in a machine-readable form. In this paper, we propose a language model adaptation framework that uses character recogni…

Cited by 0SourceScholar