← Search

Yuma Koizumi

22 accepted papers

2020

Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement

ICASSP 2020accepted

We propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-t…

Cited by 0SourceScholar
2020

Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks

ICASSP 2020accepted

Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data, several studies presented deep neural network (DNN)–based phase…

Cited by 0SourceScholar
2020

Real-Time Speech Enhancement Using Equilibriated RNN

ICASSP 2020accepted

We propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capabili…

Cited by 44SourceScholar
2020

SPIDERnet: Attention Network For One-Shot Anomaly Detection In Sounds

ICASSP 2020accepted

We propose a similarity function for one-shot anomaly detection in sounds (ADS) called SPecific anomaly IDentifiER network (SPIDERnet). In ADS systems, since overlooking an anomaly may result in serious incidents, we need to update such systems using an (often only one) overlooked anomalous sample.…

Cited by 0SourceScholar
2020

Sound Event Detection by Multitask Learning of Sound Events and Scenes with Soft Scene Labels

ICASSP 2020accepted

Sound event detection (SED) and acoustic scene classification (ASC) are major tasks in environmental sound analysis. Considering that sound events and scenes are closely related to each other, some works have addressed joint analyses of sound events and acoustic scenes based on multitask learning (M…

Cited by 0SourceScholar
2020

Sound Event Localization Based on Sound Intensity Vector Refined by Dnn-Based Denoising and Source Separation

ICASSP 2020accepted

We propose a direction-of-arrival (DOA) estimation method for Sound Event Localization and Detection (SELD). Direct estimation of DOA using a deep neural network (DNN), i.e. completely-datadriven approach, achieves high accuracy. However, there is a gap in the accuracy between DOA estimation for sin…

Cited by 0SourceScholar
2020

Speech Enhancement Using Self-Adaptation and Multi-Head Self-Attention

ICASSP 2020accepted

This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studies of deep neural network (DNN)-based speech enhancement mainly focus on buildin…

Cited by 0SourceScholar
2020

Stable Training of Dnn for Speech Enhancement Based on Perceptually-Motivated Black-Box Cost Function

ICASSP 2020accepted

Improving subjective sound quality of enhanced signals is one of the most important missions in speech enhancement. For evaluating the subjective quality, several methods related to perceptually-motivated objective sound quality assessment (OSQA) have been proposed such as PESQ (perceptual evaluatio…

Cited by 0SourceScholar
2019

A Two-class Hyper-spherical Autoencoder for Supervised Anomaly Detection

ICASSP 2019accepted

Supervised anomaly detection has been a tough problem due to its necessity of special handling of unseen anomalies. In this paper, we present a heuristic implementation of variational auto-encoder with von-Mises Fisher prior applied to a supervised anomaly detector. The closed latent space like sphe…

Cited by 0SourceScholar
2019

AdaFlow: Domain-adaptive Density Estimator with Application to Anomaly Detection and Unpaired Cross-domain Translation

ICASSP 2019accepted

We tackle unsupervised anomaly detection (UAD), a problem of detecting data that significantly differ from normal data. UAD is typically solved by using density estimation. Recently, deep neural network (DNN)-based density estimators, such as Normalizing Flows, have been attracting attention. Howeve…

Cited by 0SourceScholar
2019

Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement

ICASSP 2019accepted

We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when…

Cited by 0SourceScholar
2019

SNIPER: Few-shot Learning for Anomaly Detection to Minimize False-negative Rate with Ensured True-positive Rate

ICASSP 2019accepted

In anomaly detection systems, overlooking anomalies may result in serious incidents. Thus, when a system overlooks an anomaly, we need to update the system to never overlook the observed type of anomalies twice. There are roughly two possible approaches to solve this problem; re-training the whole s…

Cited by 0SourceScholar
2018

Complementary Set Variational Autoencoder for Supervised Anomaly Detection

ICASSP 2018accepted

Anomalies have broad patterns corresponding to their causes. In industry, anomalies are typically observed as equipment failures. Anomaly detection aims to detect such failures as anomalies. Although this is usually a binary classification task, the potential existence of unseen (unknown) failures m…

Cited by 0SourceScholar
2018

End-to-End Sound Source Enhancement Using Deep Neural Network in the Modified Discrete Cosine Transform Domain

ICASSP 2018accepted

This paper presents an end-to-end deep neural network (DNN)-based source enhancement on the basis of a time-frequency (T-F) mask processing in the modified discrete cosine transform (MDCT)-domain. To retrieve the target signal perfectly in the discrete Fourier transform (DFT)-domain, both amplitude…

Cited by 0SourceScholar
2017

DNN-based source enhancement self-optimized by reinforcement learning using sound quality measurements

ICASSP 2017accepted

We investigated whether a deep neural network (DNN)-based source enhancement function can be self-optimized by reinforcement learning (RL). The use of a DNN is a powerful approach to describing the relationship between two sets of variables and can be useful for source enhancement function design. B…

Cited by 0SourceScholar
2017

On relationships between amplitude and phase of short-time Fourier transform

ICASSP 2017accepted

The relationships between the amplitude and phase of the short-time Fourier transform (STFT) are investigated. By choosing the Gaussian window for the STFT, we reveal that the group delay and instantaneous frequency of each signal segment, both of which are derived from the phase by definition, can…

Cited by 0SourceScholar
2017

Supervised source enhancement composed of nonnegative auto-encoders and complementarity subtraction

ICASSP 2017accepted

A method for constructing deep neural networks (DNNs) for accurate supervised source enhancement is proposed. Attempts were made in previous studies to estimate the power spectral densities (PSDs) of sound sources, which are used to estimate Wiener filters for source enhancement, from the output of…

Cited by 0SourceScholar
2016

Binaural sound generation corresponding to omnidirectional video view using angular region-wise source enhancement

ICASSP 2016accepted

Web applications for watching omnidirectional video through head-mounted displays (HMDs) or smartphones have been widely distributed. The goal of this study was to generate binaural sounds corresponding to the user viewpoint. Assuming that a microphone array is used for sound recording, the enhanced…

Cited by 0SourceScholar
2016

Integrated approach of feature extraction and sound source enhancement based on maximization of mutual information

ICASSP 2016accepted

We investigated informative acoustic feature extraction based on dimension reduction for collecting target sources on a noisy sports field. Although a Wiener filter is often used for sound source enhancement, it is difficult to accurately design the Wiener filter by simply using spatial cues because…

Cited by 0SourceScholar
2016

Pinpoint extraction of distant sound source based on DNN mapping from multiple beamforming outputs to prior SNR

ICASSP 2016accepted

We propose a method for estimating the prior signal-to-noise ratio (SNR), which is used for calculating the Wiener filter for distant sound source extraction, from output signals of beamforming using statistical mapping based on the deep neural network (DNN). Since informative features to estimate t…

Cited by 17SourceScholar