← Search

Hiroshi Sato

14 accepted papers

2025

Alignment-Free Training for Transducer-based Multi-Talker ASR

ICASSP 2025accepted

Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to achieve recognition without relying on costly front-end source separation. MT-RNNT is conventionally implemented using arch…

Cited by 0SourceScholar
2025

Guided Speaker Embedding

ICASSP 2025accepted

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment…

Cited by 35SourceScholar
2024

How Does End-To-End Speech Recognition Training Impact Speech Enhancement Artifacts?

ICASSP 2024accepted

Jointly training a speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end has been investigated as a way to mitigate the influence of processing distortion generated by single-channel SE on ASR. In this paper, we investigate the effect of such joint training on the sign…

Cited by 0SourceScholar
2024

Noise-Robust Zero-Shot Text-to-Speech Synthesis Conditioned on Self-Supervised Speech-Representation Model with Adapters

ICASSP 2024accepted

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this approach suffers from degradation in speech synthesis quality when…

Cited by 0SourceScholar
2024

Talking Face Generation for Impression Conversion Considering Speech Semantics

ICASSP 2024accepted

This study investigates the talking face generation method to convert a speaker’s video to give a target impression, such as “favorable” or “considerate”. Such an impression conversion method needs to consider the input speech semantics because they affect the impression of a speaker’s video along w…

Cited by 0SourceScholar
2023

Improving Scheduled Sampling for Neural Transducer-Based ASR

ICASSP 2023accepted

The recurrent neural network-transducer (RNNT) is a promising approach for automatic speech recognition (ASR) with the introduction of a prediction network that autoregressively considers linguistic aspects. To train the autoregressive part, the ground-truth tokens are used as substitutions for the…

Cited by 0SourceScholar
2023

Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning

ICASSP 2023accepted

In this paper, we propose novel cross-lingual self-supervised speech representation learning methods that explicitly consider language information. Cross-lingual self-supervised speech representation learning has been studied to make effective use of diverse data in various languages. Previous metho…

Cited by 0SourceScholar
2022

Customer Satisfaction Estimation Using Unsupervised Representation Learning with Multi-Format Prediction Loss

ICASSP 2022accepted

We propose a new Customer Satisfaction Estimation (CSE) method that utilizes unsupervised representation learning. Though conventional methods have improved both the heuristic features and the estimation models, their performance is still insufficient as only small amounts of labeled training data c…

Cited by 0SourceScholar
2022

Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration

ICASSP 2022accepted

In this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (…

Cited by 0SourceScholar
2022

Learning to Enhance or Not: Neural Network-Based Switching of Enhanced and Observed Signals for Overlapping Speech Recognition

ICASSP 2022accepted

The combination of a deep neural network (DNN) -based speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end is a widely used approach to implement overlapping speech recognition. However, the SE front-end generates processing artifacts that can degrade the ASR performa…

Cited by 0SourceScholar
2021

Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech Recognition

ICASSP 2021accepted

Recurrent neural network-transducer (RNN-T) is promising for building time-synchronous end-to-end automatic speech recognition (ASR) systems, in part because it does not need frame-wise alignment between input features and target labels in the training step. Although training without alignment is be…

Cited by 8SourceScholar
2021

Speech Emotion Recognition Based on Listener Adaptive Models

ICASSP 2021accepted

This paper presents a novel speech emotion recognition scheme that can deal with the individuality of emotion perception. Most conventional methods directly model the majority decision of multiple listener’s perceived emotions. However, emotion perception varies with the listener, which means the co…

Cited by 0SourceScholar
2020

Distilling Attention Weights for CTC-Based ASR Systems

ICASSP 2020accepted

We present a novel training approach for connectionist temporal classification (CTC) -based automatic speech recognition (ASR) systems. CTC models are promising for building both a conventional acoustic model and an end-to-end (E2E) ASR model. However, CTC models make it difficult to capture the cor…

Cited by 0SourceScholar