← Search

Ryoichi Takashima

7 accepted papers

2023

Zero-Shot Sound Event Classification Using a Sound Attribute Vector with Global and Local Feature Learning

ICASSP 2023accepted

This paper introduces a zero-shot sound event classification (ZS-SEC) method to identify sound events that have never occurred in training data. In our previous work, we proposed a ZS-SEC method using sound attribute vectors (SAVs), where a deep neural network model infers attribute information that…

Cited by 0SourceScholar
2022

Speaker-Targeted Audio-Visual Speech Recognition Using a Hybrid CTC/Attention Model with Interference Loss

ICASSP 2022accepted

Audio-visual (AV)-automatic speech recognition (ASR) can improve speech recognition accuracy by using lip images, especially in noisy environments. The recently proposed AV Align system integrates speech and image features based on a cross-modal attention mechanism, where attention weights for visua…

Cited by 0SourceScholar
2021

High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC

ICASSP 2021accepted

This paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate…

Cited by 0SourceScholar
2020

Two-Step Acoustic Model Adaptation for Dysarthric Speech Recognition

ICASSP 2020accepted

This paper introduces a model adaptation approach for a speaker-dependent dysarthric speech recognition system. The dysarthria we focus on in this paper is caused by athetoid cerebral palsy, which causes involuntary muscle movements in those with the disease. For this reason, the dysarthric people's…

Cited by 0SourceScholar
2019

Investigation of Sequence-level Knowledge Distillation Methods for CTC Acoustic Models

ICASSP 2019accepted

This paper presents knowledge distillation (KD) methods for training connectionist temporal classification (CTC) acoustic models. In a previous study, we proposed a KD method based on the sequence-level cross-entropy, and showed that the conventional KD method based on the frame-level cross-entropy…

Cited by 0SourceScholar
2018

An Investigation of a Knowledge Distillation Method for CTC Acoustic Models

ICASSP 2018accepted

End-to-end acoustic models, such as connectionist temporal classification (CTC) and the attention model, have been studied, and their speech recognition accuracies come close to those of conventional deep neural network (DNN)-hidden Markov models. However, most high-performance end-to-end models are…

Cited by 0SourceScholar