← Search

Kenji Nagamatsu

5 accepted papers

2021

Audio-Visual Speech Enhancement Method Conditioned in the Lip Motion and Speaker-Discriminative Embeddings

ICASSP 2021accepted

We propose an audio-visual speech enhancement (AVSE) method conditioned both on the speaker’s lip motion and on speaker-discriminative embeddings. We particularly explore a method of extracting the embeddings directly from noisy audio in the AVSE setting without an enrollment procedure. We aim to im…

Cited by 0SourceScholar
2021

End-To-End Speaker Diarization as Post-Processing

ICASSP 2021accepted

This paper investigates the utilization of an end-to-end diarization model as post-processing of conventional clustering-based diarization. Clustering-based diarization methods partition frames into clusters of the number of speakers; thus, they typically cannot handle overlapping speech because eac…

Cited by 0SourceScholar
2020

Anticipating the Start of User Interaction for Service Robot in the Wild

ICRA 2020poster

A service robot is expected to provide proactive service for visitors who require its help. In contrast to passive service, e.g., providing service only after being spoken to, proactive service initiates an interaction at an early stage, e.g., talking to potential visitors who need the robot’s help…

Cited by 9SourceScholar
2019

Acoustic Modeling for Distant Multi-talker Speech Recognition with Single- and Multi-channel Branches

ICASSP 2019accepted

This paper presents a novel heterogeneous-input multi-channel acoustic model (AM) that has both single-channel and multi-channel input branches. In our proposed training pipeline, a single-channel AM is trained first, then a multi-channel AM is trained starting from the single-channel AM with a rand…

Cited by 0SourceScholar