← Search

Takuya Higuchi

13 accepted papers

2025

A Variational Framework for Improving Naturalness in Generative Spoken Language Models

ICML 2025poster

The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized for autoregressive modeling. Speech tokens derived from self-supervised models (known as semantic tokens) typically focus…

2025

Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

ICASSP 2025accepted

Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmas…

Cited by 0SourceScholar
2025

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

ICASSP 2025accepted

Iterative self-training, or iterative pseudo-labeling (IPL)—using an improved model from the current iteration to provide pseudo-labels for the next iteration—has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker re…

Cited by 0SourceScholar
2025

Towards Automatic Assessment of Self-Supervised Speech Models using Rank

ICASSP 2025accepted

This study explores using embedding rank as an unsupervised evaluation metric for general-purpose speech encoders trained via self-supervised learning (SSL). Traditionally, assessing the performance of these encoders is resource-intensive and requires labeled data from the downstream tasks. Inspired…

Cited by 0SourceScholar
2021

Dynamic Curriculum Learning via Data Parameters for Noise Robust Keyword Spotting

ICASSP 2021accepted

We propose dynamic curriculum learning via data parameters for noise robust keyword spotting. Data parameter learning has recently been introduced for image processing, where weight parameters, so-called data parameters, for target classes and instances are introduced and optimized along with model…

Cited by 0SourceScholar
2018

Dual Frequency- and Block-Permutation Alignment for Deep Learning Based Block-Online Blind Source Separation

ICASSP 2018accepted

Deep attractor networks (DANs) are a recently introduced method to blindly separate sources from spectral features of a monaural recording using bidirectional long short-term memory networks (BLSTMs). Due to the nature of BLSTMs, this is inherently not online-ready and resorting to operating on bloc…

Cited by 4SourceScholar
2018

Frame-by-Frame Closed-Form Update for Mask-Based Adaptive MVDR Beamforming

ICASSP 2018accepted

Beamforming approaches using time-frequency masks have recently been investigated and have shown promising results for noise robust automatic speech recognition (ASR) in many tasks. The time-frequency masks are estimated to compute the spatial statistics of target speech and noise signals, and then…

Cited by 0SourceScholar
2018

Optimization of Speaker-Aware Multichannel Speech Extraction with ASR Criterion

ICASSP 2018accepted

This paper addresses the problem of recognizing speech corrupted by overlapping speakers in a multichannel setting. To extract a target speaker from the mixture, we use a neural network based beamformer which uses masks estimated by a neural network to compute statistically optimal spatial filters.…

Cited by 0SourceScholar
2017

Deep mixture density network for statistical model-based feature enhancement

ICASSP 2017accepted

We propose a novel framework designed to extend conventional deep neural network (DNN)-based feature enhancement approaches. In general, the conventional DNN-based feature enhancement framework aims to map input noisy observation to clean speech or a binary/ soft mask in a deterministic way, assumin…

Cited by 13SourceScholar
2017

Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming

ICASSP 2017accepted

Recently, time-frequency mask-based beamforming has been extensively studied as the frontend of deep neural network (DNN) based automatic speech recognition (ASR) in noisy environments. Two mask estimation approaches have been separately developed for this beamforming method, namely the the DNN-base…

Cited by 0SourceScholar
2017

Unsupervised utterance-wise beamformer estimation with speech recognition-level criterion

ICASSP 2017accepted

In this paper, we perform beamforming with a speech recognition-level criterion. A beamformer is usually designed by optimizing signal-level criteria, e.g., by minimizing the beamformer output covariance or by maximizing the signal-to-noise ratio (SNR). Such signal-level criteria do not always guara…

Cited by 0SourceScholar
2016

Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise

ICASSP 2016accepted

This paper considers acoustic beamforming for noise robust automatic speech recognition (ASR). A beamformer attenuates background noise by enhancing sound components coming from a direction specified by a steering vector. Hence, accurate steering vector estimation is paramount for successful noise r…

Cited by 0SourceScholar
2016

Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognition

ICASSP 2016accepted

This paper addresses a minimum variance distortionless response (MVDR) beamforming based speech enhancement approach for meeting speech recognition. In a meeting situation, speaker overlaps and noise signals are not negligible. To handle these issues, we employ MVDR beamforming, where accurate estim…

Cited by 0SourceScholar