← Search

Ryo Aihara

6 accepted papers

2026

FLEXIO: FLEXIBLE SINGLE- AND MULTI-CHANNEL SPEECH SEPARATION AND ENHANCEMENT

ICASSP 2026oral

Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speak…

Cited by 0SourcePDFScholar
2022

Speaker-Targeted Audio-Visual Speech Recognition Using a Hybrid CTC/Attention Model with Interference Loss

ICASSP 2022accepted

Audio-visual (AV)-automatic speech recognition (ASR) can improve speech recognition accuracy by using lip images, especially in noisy environments. The recently proposed AV Align system integrates speech and image features based on a cross-modal attention mechanism, where attention weights for visua…

Cited by 0SourceScholar
2019

Teacher-student Deep Clustering for Low-delay Single Channel Speech Separation

ICASSP 2019accepted

The recently-proposed deep clustering algorithm introduced significant advances in monaural speaker-independent multi-speaker speech separation. Deep clustering operates on magnitude spectro-grams using bidirectional recurrent networks and K-means clustering, both of which require offline operation,…

Cited by 0SourceScholar
2016

Semi-non-negative matrix factorization using alternating direction method of multipliers for voice conversion

ICASSP 2016accepted

Voice conversion (VC) is being widely researched in the field of speech processing because of increased interest in using such processing in applications such as personalized Text-To-Speech systems. A VC method using Non-negative Matrix Factorization (NMF) has been researched because of its natural…

Cited by 0SourceScholar
2015

Activity-mapping non-negative matrix factorization for exemplar-based voice conversion

ICASSP 2015accepted

Voice conversion (VC) is being widely researched in the field of speech processing because of increased interest in using such processing in applications such as personalized Text-To-Speech systems. We present in this paper an exemplar-based VC method us- ing Non-negative Matrix Factorization (NMF),…

Cited by 0SourceScholar