← Search

Takahiro Shinozaki

14 accepted papers

2025

Deep Generic Representations for Domain-Generalized Anomalous Sound Detection

ICASSP 2025accepted

Developing a reliable anomalous sound detection (ASD) system requires robustness to noise, adaptation to domain shifts, and effective performance with limited training data. Current leading methods rely on extensive labeled data for each target machine type to train feature extractors using Outlier-…

Cited by 0SourceScholar
2024

Self-Supervised Speaker Verification with Adaptive Threshold and Hierarchical Training

ICASSP 2024accepted

In self-supervised speaker verification, the quality of generated pseudo labels becomes a bottleneck for the performance. This work introduces a dynamic threshold within the iterative DIstillation with NO labels (DINO) framework. We employ a Gaussian Mixture Model (GMM) to model the loss distributio…

Cited by 0SourceScholar
2023

Continuous Action Space-Based Spoken Language Acquisition Agent Using Residual Sentence Embedding and Transformer Decoder

ICASSP 2023accepted

Studies on spoken language acquisition agents aim to understand the mechanism of human language learning and to realize it on computers. Existing open vocabulary agents first perform unsupervised word learning from speech signals to construct a word dictionary as a discrete action space and then con…

Cited by 0SourceScholar
2023

FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning

ICLR 2023poster

Semi-supervised Learning (SSL) has witnessed great success owing to the impressive performances brought by various methods based on pseudo labeling and consistency regularization. However, we argue that existing methods might fail to utilize the unlabeled data more effectively since they either use…

2023

Multi-Domain Dialogue State Tracking with Disentangled Domain-Slot Attention

ACL 2023findings

As the core of task-oriented dialogue systems, dialogue state tracking (DST) is designed to track the dialogue state through the conversation between users and systems. Multi-domain DST has been an important challenge in which the dialogue states across multiple domains need to consider. In recent m…

Cited by 5SourcePDFScholar
2022

Exploiting Unlabeled Data for Target-Oriented Opinion Words Extraction

COLING 2022main

Target-oriented Opinion Words Extraction (TOWE) is a fine-grained sentiment analysis task that aims to extract the corresponding opinion words of a given opinion target from the sentence. Recently, deep learning approaches have made remarkable progress on this task. Nevertheless, the TOWE task still…

2022

Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration

ICASSP 2022accepted

In this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (…

Cited by 0SourceScholar
2022

USB: A Unified Semi-supervised Learning Benchmark for Classification

NeurIPS 2022accept

Semi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural netw…

2021

FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling

NeurIPS 2021poster

The recently proposed FixMatch achieved state-of-the-art results on most semi-supervised learning (SSL) benchmarks. However, like other modern SSL algorithms, FixMatch uses a pre-defined constant threshold for all classes to select unlabeled data that contribute to the training, thus failing to cons…

2021

Meta-Adapter: Efficient Cross-Lingual Adaptation With Meta-Learning

ICASSP 2021accepted

Transfer learning from a multilingual model has shown favorable results on low-resource automatic speech recognition (ASR). However, full-model fine-tuning generates a separate model for every target language and is not suitable for deploying and maintaining in production. The key challenge lies in…

Cited by 0SourceScholar
2020

Spoken Language Acquisition Based on Reinforcement Learning and Word Unit Segmentation

ICASSP 2020accepted

The process of spoken-language acquisition has been one of the topics of greatest interest to linguists for decades. By uti-lizing modern machine learning techniques, we simulated this process on computers, which helps to understand it and develop new possibilities of applying this concept on intell…

Cited by 0SourceScholar
2019

Effective and Stable Neuron Model Optimization Based on Aggregated CMA-ES

ICASSP 2019accepted

Computer simulations have facilitated our understanding of the dynamic behavior of the brain and the effect of the medical treatment such as deep brain stimulation. For improving the simulation model, it is essential to develop a method for optimizing parameters of a neuron model from available expe…

Cited by 0SourceScholar
2018

Reinforcement Learning of Speech Recognition System Based on Policy Gradient and Hypothesis Selection

ICASSP 2018accepted

Automatic speech recognition (ASR) systems have achieved high recognition performance for several tasks. However, the performance of such systems is dependent on the tremendously costly development work of preparing vast amounts of task-matched transcribed speech data for supervised training. The ke…

Cited by 0SourceScholar