← Search

Atsushi Ando

12 accepted papers

2025

Guided Speaker Embedding

ICASSP 2025accepted

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment…

Cited by 35SourceScholar
2025

Mamba-based Segmentation Model for Speaker Diarization

ICASSP 2025accepted

Mamba is a newly proposed architecture that behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are…

Cited by 14SourceScholar
2025

Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation

ICASSP 2025accepted

This paper proposes a speaker counting scheme using multichannel microphones for end-to-end neural diarization with a vector clustering (EEND-VC) speaker diarization pipeline. The EEND-VC-based system estimates the number of speakers by clustering speaker embeddings from small chunks. However, conve…

Cited by 0SourceScholar
2025

Speech Emotion Recognition Based on Large-Scale Automatic Speech Recognizer

ICASSP 2025accepted

This paper proposes a novel speech emotion recognition (SER) method that fully leverages the architecture of Whisper, a large-scale automatic speech recognition (ASR) model. The conventional SER models using a pre-trained speech encoder may fail to capture linguistic content since their decoders are…

Cited by 0SourceScholar
2024

NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization

ICASSP 2024accepted

This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)based dereverberation as a front end, and separately applies end-to-end neural diarization with vector clustering…

Cited by 0SourceScholar
2023

Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness Tradeoff

ICCV 2023poster

This paper addresses the tradeoff between standard accuracy on clean examples and robustness against adversarial examples in deep neural networks (DNNs). Although adversarial training (AT) improves robustness, it degrades the standard accuracy, thus yielding the tradeoff. To mitigate this tradeoff…

Cited by 7PDFScholar
2022

Customer Satisfaction Estimation Using Unsupervised Representation Learning with Multi-Format Prediction Loss

ICASSP 2022accepted

We propose a new Customer Satisfaction Estimation (CSE) method that utilizes unsupervised representation learning. Though conventional methods have improved both the heuristic features and the estimation models, their performance is still insufficient as only small amounts of labeled training data c…

Cited by 0SourceScholar
2022

Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration

ICASSP 2022accepted

In this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (…

Cited by 0SourceScholar
2021

Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech Recognition

ICASSP 2021accepted

Recurrent neural network-transducer (RNN-T) is promising for building time-synchronous end-to-end automatic speech recognition (ASR) systems, in part because it does not need frame-wise alignment between input features and target labels in the training step. Although training without alignment is be…

Cited by 8SourceScholar
2021

Speech Emotion Recognition Based on Listener Adaptive Models

ICASSP 2021accepted

This paper presents a novel speech emotion recognition scheme that can deal with the individuality of emotion perception. Most conventional methods directly model the majority decision of multiple listener’s perceived emotions. However, emotion perception varies with the listener, which means the co…

Cited by 0SourceScholar
2020

Sequence-Level Consistency Training for Semi-Supervised End-to-End Automatic Speech Recognition

ICASSP 2020accepted

This paper presents a novel semi-supervised end-to-end automatic speech recognition (ASR) method that employs consistency training with the use of unlabeled data. In consistency training, unlabeled data can be utilized for constraining a model such that it becomes invariant to small deformation. In…

Cited by 0SourceScholar
2018

Soft-Target Training with Ambiguous Emotional Utterances for DNN-Based Speech Emotion Classification

ICASSP 2018accepted

This paper presents a novel emotion classification method for natural speech. One of the problems in the state-of-the-art method based on Deep Neural Network (DNN) is the paucity of the training data compared to model complexity. To solve this problem, this paper utilizes the ambiguous emotional utt…

Cited by 0SourceScholar