← Search

Naohiro Tawara

18 accepted papers

2025

Guided Speaker Embedding

ICASSP 2025accepted

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment…

Cited by 35SourceScholar
2025

Mamba-based Segmentation Model for Speaker Diarization

ICASSP 2025accepted

Mamba is a newly proposed architecture that behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are…

Cited by 14SourceScholar
2025

Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation

ICASSP 2025accepted

This paper proposes a speaker counting scheme using multichannel microphones for end-to-end neural diarization with a vector clustering (EEND-VC) speaker diarization pipeline. The EEND-VC-based system estimates the number of speakers by clustering speaker embeddings from small chunks. However, conve…

Cited by 0SourceScholar
2025

SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model

ICASSP 2025accepted

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the sa…

Cited by 0SourceScholar
2024

Discriminative Training of VBx Diarization

ICASSP 2024accepted

Bayesian HMM clustering of x-vector sequences (VBx) has become a widely adopted diarization baseline model in publications and challenges. It uses an HMM to model speaker turns, a generatively trained probabilistic linear discriminant analysis (PLDA) for speaker distribution modeling, and Bayesian i…

Cited by 0SourceScholar
2024

NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization

ICASSP 2024accepted

This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)based dereverberation as a front end, and separately applies end-to-end neural diarization with vector clustering…

Cited by 0SourceScholar
2023

Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition

ICASSP 2023accepted

We propose a new shallow fusion (SF) method to exploit an external backward language model (BLM) for end-to-end automatic speech recognition (ASR). The BLM has complementary characteristics with a forward language model (FLM), and the effectiveness of their combination has been confirmed by rescorin…

Cited by 0SourceScholar
2022

Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models

ICASSP 2022accepted

We investigate the effectiveness of using a large ensemble of advanced neural language models (NLMs) for lattice rescoring on automatic speech recognition (ASR) hypotheses. Previous studies have reported the effectiveness of combining a small number of NLMs. In contrast, in this study, we combine up…

Cited by 0SourceScholar
2021

Age-VOX-Celeb: Multi-Modal Corpus for Facial and Speech Estimation

ICASSP 2021accepted

Estimating a speaker’s age from their speech is more challenging than age estimation from their face because of insufficiently available public corpora. To tackle this problem, we construct a new audio-visual age corpus named AgeVoxCeleb by annotating age labels to VoxCeleb2 videos. AgeVoxCeleb is t…

Cited by 0SourceScholar
2021

BLSTM-Based Confidence Estimation for End-to-End Speech Recognition

ICASSP 2021accepted

Confidence estimation, in which we estimate the reliability of each recognized token (e.g., word, sub-word, and character) in automatic speech recognition (ASR) hypotheses and detect incorrectly recognized tokens, is an important function for developing ASR applications. In this study, we perform co…

Cited by 0SourceScholar
2021

Integrating End-to-End Neural and Clustering-Based Diarization: Getting the Best of Both Worlds

ICASSP 2021accepted

Recent diarization technologies can be categorized into two approaches, i.e., clustering and end-to-end neural approaches, which have different pros and cons. The clustering-based approaches assign speaker labels to speech regions by clustering speaker embeddings such as x-vectors. While it can be s…

Cited by 0SourceScholar
2020

Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances

ICASSP 2020accepted

This paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and tes…

Cited by 0SourceScholar
2020

Improving Speaker Discrimination of Target Speech Extraction With Time-Domain Speakerbeam

ICASSP 2020accepted

Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation utterance of the target speaker to extract his/her voice characteristics that are…

Cited by 152SourceScholar
2020

Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster Information

ICASSP 2020accepted

This paper proposes a general post-processing method for improving speaker-attribute estimation. Estimating speaker-specific attributes such as age and gender is an important task with a wide range of applications. While the recent proposed deep neural network-based end-to-end approach achieves high…

Cited by 0SourceScholar
2019

Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware Training

ICASSP 2019accepted

An adversarial denoising autoencoder (ADAE) with noise-aware training is proposed and successfully applied to post-filtering for linear noise reduction. The ADAE is effective for attenuating interference sounds, however, it is difficult to learn to handle its various unexpected harmful effects (e.g.…

Cited by 2SourceScholar
2018

Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations

ICASSP 2018accepted

Training recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to…

Cited by 0SourceScholar
2018

Speaker Invariant Feature Extraction for Zero-Resource Languages with Adversarial Learning

ICASSP 2018accepted

We introduce a novel type of representation learning to obtain a speaker invariant feature for zero-resource languages. Speaker adaptation is an important technique to build a robust acoustic model. For a zero-resource language, however, conventional model-dependent speaker adaptation methods such a…

Cited by 0SourceScholar
2015

A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditions

ICASSP 2015accepted

The present paper dealt with speaker clustering for speech corrupted by noise. In general, the performance of speaker clustering significantly depends on how well the similarities between speech utterances can be measured. The recently proposed i-vector-based cosine similarity has yielded the state-…

Cited by 8SourceScholar