← Search

Shota Horiguchi

11 accepted papers

2025

Alignment-Free Training for Transducer-based Multi-Talker ASR

ICASSP 2025accepted

Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to achieve recognition without relying on costly front-end source separation. MT-RNNT is conventionally implemented using arch…

Cited by 0SourceScholar
2025

Guided Speaker Embedding

ICASSP 2025accepted

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment…

Cited by 0SourceScholar
2025

Mamba-based Segmentation Model for Speaker Diarization

ICASSP 2025accepted

Mamba is a newly proposed architecture that behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are…

Cited by 0SourceScholar
2025

Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation

ICASSP 2025accepted

This paper proposes a speaker counting scheme using multichannel microphones for end-to-end neural diarization with a vector clustering (EEND-VC) speaker diarization pipeline. The EEND-VC-based system estimates the number of speakers by clustering speaker embeddings from small chunks. However, conve…

Cited by 0SourceScholar
2024

Streaming Active Learning for Regression Problems Using Regression via Classification

ICASSP 2024accepted

One of the challenges in deploying a machine learning model is that the model’s performance degrades as the operating environment changes. To maintain the performance, streaming active learning is used, in which the model is retrained by adding a newly annotated sample to the training dataset if the…

Cited by 0SourceScholar
2022

Environmental Sound Extraction Using Onomatopoeic Words

ICASSP 2022accepted

An onomatopoeic word, which is a character sequence that phonetically imitates a sound, is effective in expressing characteristics of sound such as duration, pitch, and timbre. We propose an environmental-sound-extraction method using onomatopoeic words to specify the target sound to be extracted. B…

Cited by 0SourceScholar
2022

Multi-Channel End-To-End Neural Diarization with Distributed Microphones

ICASSP 2022accepted

Recent progress on end-to-end neural diarization (EEND) has en-abled overlap-aware speaker diarization with a single neural net-work. This paper proposes to enhance EEND by using multi-channel signals from distributed microphones. We replace Transformer en-coders in EEND with two types of encoders t…

Cited by 0SourceScholar
2022

Rethinking Fano’s Inequality in Ensemble Learning

ICML 2022spotlight

We propose a fundamental theory on ensemble learning that evaluates a given ensemble system by a well-grounded set of metrics. Previous studies used a variant of Fano’s inequality of information theory and derived a lower bound of the classification error rate on the basis of the accuracy and divers…

2021

End-To-End Speaker Diarization as Post-Processing

ICASSP 2021accepted

This paper investigates the utilization of an end-to-end diarization model as post-processing of conventional clustering-based diarization. Clustering-based diarization methods partition frames into clusters of the number of speakers; thus, they typically cannot handle overlapping speech because eac…

Cited by 0SourceScholar
2020

Anticipating the Start of User Interaction for Service Robot in the Wild

ICRA 2020poster

A service robot is expected to provide proactive service for visitors who require its help. In contrast to passive service, e.g., providing service only after being spoken to, proactive service initiates an interaction at an early stage, e.g., talking to potential visitors who need the robot’s help…

Cited by 9SourceScholar
2019

Acoustic Modeling for Distant Multi-talker Speech Recognition with Single- and Multi-channel Branches

ICASSP 2019accepted

This paper presents a novel heterogeneous-input multi-channel acoustic model (AM) that has both single-channel and multi-channel input branches. In our proposed training pipeline, a single-channel AM is trained first, then a multi-channel AM is trained starting from the single-channel AM with a rand…

Cited by 0SourceScholar