← Search

Daisuke Saito

6 accepted papers

2026

ADVANCED MODELING OF INTERLANGUAGE SPEECH INTELLIGIBILITY BENEFIT WITH L1-L2 MULTI-TASK LEARNING USING DIFFERENTIABLE K-MEANS FOR ACCENT-ROBUST DISCRETE TOKEN-BASED ASR

ICASSP 2026poster

Building ASR systems robust to foreign-accented speech is an important challenge in today's globalized world. A prior study explored the way to enhance the performance of phonetic token-based ASR on accented speech by reproducing the phenomenon known as interlanguage speech intelligibility benefit (…

Cited by 0SourcePDFScholar
2024

Do Learned Speech Symbols Follow Zipf's Law?

ICASSP 2024accepted

In this study, we investigate whether speech symbols, learned through deep learning, follow Zipf’s law, akin to natural language symbols. Zipf’s law is an empirical law that delineates the frequency distribution of words, forming fundamentals for statistical analysis in natural language processing.…

Cited by 0SourceScholar
2023

Multiple Acoustic Features Speech Emotion Recognition Using Cross-Attention Transformer

ICASSP 2023accepted

Speech emotion recognition (SER) is a challenging task whose performance heavily relies on suitable affect-salient representations. Recently, transformer has exhibited outstanding qualities in learning relevant representations associated with this task. However, a normal transformer is only able to…

Cited by 0SourceScholar
2016

Divergence estimation based on deep neural networks and its use for language identification

ICASSP 2016accepted

In this paper, we propose a method to estimate statistical divergence between probability distributions by a DNN-based discriminative approach and its use for language identification tasks. Since statistical divergence is generally defined as a functional of two probability density functions, these…

Cited by 0SourceScholar
2015

SAS: A speaker verification spoofing database containing diverse attacks

ICASSP 2015accepted

This paper presents the first version of a speaker verification spoofing and anti-spoofing database, named SAS corpus. The corpus includes nine spoofing techniques, two of which are speech synthesis, and seven are voice conversion. We design two protocols, one for standard speaker verification evalu…

Cited by 0SourceScholar