← Search

Tatsuya Komatsu

21 accepted papers

2025

Aligned Contrastive Learning for Text-to-Music Retrieval

ICASSP 2025accepted

This paper proposes aligned contrastive learning for text-to-music retrieval. The proposed method introduces a new similarity measure, 'aligned similarity', which captures the frame-level and token-level correspondence within text and audio sequences. Unlike traditional approaches that aggregate seq…

Cited by 0SourceScholar
2025

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

ICASSP 2025accepted

Current audio-visual representation learning can capture rough object categories (e.g., "animals" and "instruments"), but it lacks the ability to recognize fine-grained details, such as specific categories like "dogs" and "flutes" within animals and instruments. To address this issue, we introduce D…

Cited by 0SourceScholar
2024

Keep Decoding Parallel With Effective Knowledge Distillation From Language Models To End-To-End Speech Recognisers

ICASSP 2024accepted

This study presents a novel approach for knowledge distillation (KD) from a BERT teacher model to an automatic speech recognition (ASR) model using intermediate layers. To distil the teacher’s knowledge, we use an attention decoder that learns from BERT’s token probabilities. Our method shows that l…

Cited by 0SourceScholar
2024

Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection

EMNLP 2024system demonstrations

We propose Lighthouse, a user-friendly library for reproducible video moment retrieval and highlight detection (MR-HD). Although researchers proposed various MR-HD approaches, the research community holds two main issues. The first is a lack of comprehensive and reproducible experiments across vario…

2024

PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions

ICASSP 2024accepted

We propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions. To control speaker identity within the prompt-based TTS framework, we introduce the concept of speaker prompt, which describes voice characteris…

Cited by 0SourceScholar
2023

Neural Diarization with Non-Autoregressive Intermediate Attractors

ICASSP 2023accepted

End-to-end neural diarization (EEND) with encoder-decoder-based attractors (EDA) is a promising method to handle the whole speaker diarization problem simultaneously with a single neural network. While the EEND model can produce all frame-level speaker labels simultaneously, it disregards output lab…

Cited by 14SourceScholar
2022

Self-Supervised Learning Method Using Multiple Sampling Strategies for General-Purpose Audio Representation

ICASSP 2022accepted

We propose a self-supervised learning method using multiple sampling strategies to obtain general-purpose audio representation. Multiple sampling strategies are used in the proposed method to construct contrastive losses from different perspectives and learn representations based on them. In this st…

Cited by 0SourceScholar
2021

Disentangled Speaker and Language Representations Using Mutual Information Minimization and Domain Adaptation for Cross-Lingual TTS

ICASSP 2021accepted

We propose a method for obtaining disentangled speaker and language representations via mutual information minimization and domain adaptation for cross-lingual text-to-speech (TTS) synthesis. The proposed method extracts speaker and language embeddings from acoustic features by a speaker encoder and…

Cited by 0SourceScholar
2020

Consistency-Aware Multi-Channel Speech Enhancement Using Deep Neural Networks

ICASSP 2020accepted

This paper proposes a deep neural network (DNN)–based multichannel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often conducted in the time-frequency (T-F) domain because spatial filterin…

Cited by 0SourceScholar
2020

Scene-Dependent Acoustic Event Detection with Scene Conditioning and Fake-Scene-Conditioned Loss

ICASSP 2020accepted

In this paper, we propose scene-dependent acoustic event detection (AED) with scene conditioning and fake-scene-conditioned loss. The proposed method employs a multitask network, that has not only AED part but also acoustic scene classification (ASC). The scenes predicted by ASC are employed as an a…

Cited by 0SourceScholar
2020

Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss Function

ICASSP 2020accepted

In this paper, we propose a multi-channel speech source separation method with a deep neural network (DNN) which is trained under the condition that no clean signal is available. As an alternative to a clean signal, the proposed method adopts an estimated speech signal by an unsupervised speech sour…

Cited by 0SourceScholar
2020

Weakly-Supervised Sound Event Detection with Self-Attention

ICASSP 2020accepted

In this paper, we propose a novel sound event detection (SED) method that incorporates a self-attention mechanism of the Transformer for a weakly-supervised learning scenario. The proposed method utilizes the Transformer encoder, which consists of multiple self-attention modules, allowing to take bo…

Cited by 0SourceScholar
2019

Bayesian Non-parametric Multi-source Modelling Based Determined Blind Source Separation

ICASSP 2019accepted

This paper proposes a determined blind source separation method using Bayesian non-parametric modelling of sources. Conventionally source signals are separated from a given set of mixture signals by modelling them using non-negative matrix factorization (NMF). However in NMF, a latent variable signi…

Cited by 0SourceScholar
2019

Scene-dependent Anomalous Acoustic-event Detection Based on Conditional Wavenet and I-vector

ICASSP 2019accepted

This paper proposes a scene-dependent anomalous acoustic-event detection based on conditional WaveNet and i-vector. The WaveNet builds normal acoustic event models by exhaustive learning of time-domain signals in the public space to provide scene-independent anomaly detection. I-vectors are used as…

Cited by 0SourceScholar
2016

Acoustic event detection based on non-negative matrix factorization with mixtures of local dictionaries and activation aggregation

ICASSP 2016accepted

This paper proposes a new non-negative matrix factorization (NMF) based acoustic event detection (AED) method with mixtures of local dictionaries (MLD) and activation aggregation. One of the key problems of conventional NMF-based methods is instability of activations due to redundancy of a region sp…

Cited by 0SourceScholar