← Search

Tetsuya Takiguchi

10 accepted papers

2023

Zero-Shot Sound Event Classification Using a Sound Attribute Vector with Global and Local Feature Learning

ICASSP 2023accepted

This paper introduces a zero-shot sound event classification (ZS-SEC) method to identify sound events that have never occurred in training data. In our previous work, we proposed a ZS-SEC method using sound attribute vectors (SAVs), where a deep neural network model infers attribute information that…

Cited by 0SourceScholar
2022

Speaker-Targeted Audio-Visual Speech Recognition Using a Hybrid CTC/Attention Model with Interference Loss

ICASSP 2022accepted

Audio-visual (AV)-automatic speech recognition (ASR) can improve speech recognition accuracy by using lip images, especially in noisy environments. The recently proposed AV Align system integrates speech and image features based on a cross-modal attention mechanism, where attention weights for visua…

Cited by 0SourceScholar
2021

High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC

ICASSP 2021accepted

This paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate…

Cited by 0SourceScholar
2020

Two-Step Acoustic Model Adaptation for Dysarthric Speech Recognition

ICASSP 2020accepted

This paper introduces a model adaptation approach for a speaker-dependent dysarthric speech recognition system. The dysarthria we focus on in this paper is caused by athetoid cerebral palsy, which causes involuntary muscle movements in those with the disease. For this reason, the dysarthric people's…

Cited by 0SourceScholar
2018

Parallel-Data-Free Dictionary Learning for Voice Conversion Using Non-Negative Tucker Decomposition

ICASSP 2018accepted

Voice conversion (VC) is a technique where only speaker-specific information in source speech is converted while preserving the associated phonological information. Nonnegative Matrix Factorization (NMF)-based VC has been researched because of the natural-sounding voice it produces compared with con…

Cited by 0SourceScholar
2016

Modeling deep bidirectional relationships for image classification and generation

ICASSP 2016accepted

This paper presents a novel probabilistic model that represents a joint probability of two visible variables with a deep architecture, called a deep relational model (DRM). The model stacks several layers from one visible layer on to another visible layer, sandwiching hidden layers between them. As…

Cited by 0SourceScholar
2016

Semi-non-negative matrix factorization using alternating direction method of multipliers for voice conversion

ICASSP 2016accepted

Voice conversion (VC) is being widely researched in the field of speech processing because of increased interest in using such processing in applications such as personalized Text-To-Speech systems. A VC method using Non-negative Matrix Factorization (NMF) has been researched because of its natural…

Cited by 0SourceScholar
2015

Activity-mapping non-negative matrix factorization for exemplar-based voice conversion

ICASSP 2015accepted

Voice conversion (VC) is being widely researched in the field of speech processing because of increased interest in using such processing in applications such as personalized Text-To-Speech systems. We present in this paper an exemplar-based VC method us- ing Non-negative Matrix Factorization (NMF),…

Cited by 0SourceScholar