← Search

Daiki Takeuchi

9 accepted papers

2025

Collision-less and Balanced Sampling for Language-Queried Audio Source Separation

ICASSP 2025accepted

Language-queried audio source separation (LASS) is an emerging research field that has recently received increasing attention. This task aims to isolate individual sources from a mixture of signals using natural language descriptions, enabling applications in various areas such as automatic audio ed…

Cited by 0SourceScholar
2024

Unrestricted Global Phase Bias-Aware Single-Channel Speech Enhancement with Conformer-Based Metric Gan

ICASSP 2024accepted

With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally outstanding. However, enhancing the phase spectrum using neural networks is often…

Cited by 0SourceScholar
2023

Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input

ICASSP 2023accepted

Masked Autoencoders is a simple yet powerful self-supervised learning method. However, it learns representations indirectly by reconstructing masked input patches. Several methods learn representations directly by predicting representations of masked patches; however, we think using all patches to e…

Cited by 0SourceScholar
2020

Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement

ICASSP 2020accepted

We propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-t…

Cited by 0SourceScholar
2020

Real-Time Speech Enhancement Using Equilibriated RNN

ICASSP 2020accepted

We propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capabili…

Cited by 44SourceScholar
2020

Speech Enhancement Using Self-Adaptation and Multi-Head Self-Attention

ICASSP 2020accepted

This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studies of deep neural network (DNN)-based speech enhancement mainly focus on buildin…

Cited by 0SourceScholar
2019

Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement

ICASSP 2019accepted

We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when…

Cited by 0SourceScholar
2018

Parametric Approximation of Piano Sound Based on Kautz Model with Sparse Linear Prediction

ICASSP 2018accepted

The piano is one of the most popular and attractive musical instruments that leads to a lot of research on it. To synthesize the piano sound in a computer, many modeling methods have been proposed from full physical models to approximated models. The focus of this paper is on the latter, approximati…

Cited by 0SourceScholar
2018

Realizing Directional Sound Source in FDTD Method by Estimating Initial Value

ICASSP 2018accepted

Wave-based acoustic simulation methods are studied actively for predicting acoustical phenomena. Finite-difference time-domain (FDTD) method is one of the most popular methods owing to its straightforwardness of calculating an impulse response. In an FDTD simulation, an omnidirectional sound source…

Cited by 1SourceScholar