← Search

Yasuhiro Oikawa

28 accepted papers

2023

Improving Phase-Vocoder-Based Time Stretching by Time-Directional Spectrogram Squeezing

ICASSP 2023accepted

Time stretching of music signals has a crucial problem, i.e., smearing of percussive sounds. Some time stretching algorithms have addressed this problem by detecting percussive components and manipulating them differently from the other components. However, conventional methods cause artifacts. In t…

Cited by 0SourceScholar
2023

UPGLADE: Unplugged Plug-and-Play Audio Declipper Based on Consensus Equilibrium of DNN and Sparse Optimization

ICASSP 2023accepted

In this paper, we propose a novel audio declipping method that fuses sparse-optimization-based and deep neural network (DNN)– based methods. The two methods have contrasting characteristics, depending on clipping level. Sparse-optimization-based audio de-clipping can preserve reliable samples, being…

Cited by 0SourceScholar
2022

APPLADE: Adjustable Plug-and-Play Audio Declipper Combining DNN with Sparse Optimization

ICASSP 2022accepted

In this paper, we propose an audio declipping method that takes advantages of both sparse optimization and deep learning. Since sparsity-based audio declipping methods have been developed upon constrained optimization, they are adjustable and well-studied in theory. However, they always uniformly pr…

Cited by 0SourceScholar
2022

Acoustic Application of Phase Reconstruction Algorithms in Optics

ICASSP 2022accepted

Phase reconstruction from amplitude spectrograms has attracted attention in recent acoustics because of its potential applications in speech synthesis and enhancement. The most well-known algorithm in acoustics is based on alternating projection and called Griffin– Lim algorithm (GLA). At the same t…

Cited by 0SourceScholar
2022

Harmonic and Percussive Sound Separation Based on Mixed Partial Derivative of Phase Spectrogram

ICASSP 2022accepted

Harmonic and percussive sound separation (HPSS) is a widely applied pre-processing tool that extracts distinct (harmonic and percussive) components of a signal. In the previous methods, HPSS has been performed based on the structural properties of magnitude (or power) spectrograms. However, such app…

Cited by 0SourceScholar
2022

Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around Head

ICASSP 2022accepted

Sound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with the sound signals observed by a microphone array. Although ambisonic microphones are popular in the literature of SELD, t…

Cited by 0SourceScholar
2020

Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement

ICASSP 2020accepted

We propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-t…

Cited by 0SourceScholar
2020

Maximally Energy-Concentrated Differential Window for Phase-Aware Signal Processing Using Instantaneous Frequency

ICASSP 2020accepted

The short-time Fourier transform (STFT) is widely employed in non-stationary signal analysis, whose property depends on window functions. Instantaneous frequency in STFT, the time-derivative of phase, is recently applied to many applications including spectrogram reassignment. The computation of ins…

Cited by 0SourceScholar
2020

Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks

ICASSP 2020accepted

Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data, several studies presented deep neural network (DNN)–based phase…

Cited by 0SourceScholar
2020

Real-Time Speech Enhancement Using Equilibriated RNN

ICASSP 2020accepted

We propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capabili…

Cited by 44SourceScholar
2020

Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling

IROS 2020poster

Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living environments, labeling all sounding objects is impossible in practice. This calls f…

Cited by 21SourceScholar
2019

Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement

ICASSP 2019accepted

We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when…

Cited by 0SourceScholar
2019

Guided-spatio-temporal Filtering for Extracting Sound from Optically Measured Images Containing Occluding Objects

ICASSP 2019accepted

Recent development of optical interferometry enables us to measure sound without placing any device inside the sound field. In particular, parallel phase-shifting interferometry (PPSI) has realized advanced measurement of refractive index of air. Its novel application investigated very recently is s…

Cited by 0SourceScholar
2019

Low-rankness of Complex-valued Spectrogram and Its Application to Phase-aware Audio Processing

ICASSP 2019accepted

Low-rankness of amplitude spectrograms has been effectively utilized in audio signal processing methods including non-negative matrix factorization. However, such methods have a fundamental limitation owing to their amplitude-only treatment where the phase of the observed signal is utilized for resy…

Cited by 0SourceScholar
2019

Phase-aware Harmonic/percussive Source Separation via Convex Optimization

ICASSP 2019accepted

Decomposition of an audio mixture into harmonic and percussive components, namely harmonic/percussive source separation (HPSS), is a useful pre-processing tool for many audio applications. Popular approaches to HPSS exploit the distinctive source-specific structures of power spectrograms. However, s…

Cited by 0SourceScholar
2018

Individual Difference of Ultrasonic Transducers for Parametric Array Loudspeaker

ICASSP 2018accepted

A parametric array loudspeaker (PAL) consists of a lot of ultrasonic transducers in most cases and is driven by an ultrasonic which is modulated by audible sound. Because each ultrasonic transducer has each difference resonant frequency, there is the individual difference in ultrasonic transducers o…

Cited by 0SourceScholar
2018

Modal Decomposition of Musical Instrument Sound Via Alternating Direction Method of Multipliers

ICASSP 2018accepted

For a musical instrument sound containing partials, or modes, the behavior of modes around the attack time is particularly important. However, accurately decomposing it around the attack time is not an easy task, especially when the onset is sharp. This is because spectra of the modes are peaky whil…

Cited by 0SourceScholar
2018

Parametric Approximation of Piano Sound Based on Kautz Model with Sparse Linear Prediction

ICASSP 2018accepted

The piano is one of the most popular and attractive musical instruments that leads to a lot of research on it. To synthesize the piano sound in a computer, many modeling methods have been proposed from full physical models to approximated models. The focus of this paper is on the latter, approximati…

Cited by 0SourceScholar
2018

Realizing Directional Sound Source in FDTD Method by Estimating Initial Value

ICASSP 2018accepted

Wave-based acoustic simulation methods are studied actively for predicting acoustical phenomena. Finite-difference time-domain (FDTD) method is one of the most popular methods owing to its straightforwardness of calculating an impulse response. In an FDTD simulation, an omnidirectional sound source…

Cited by 1SourceScholar
2017

Coherence-adjusted monopole dictionary and convex clustering for 3D localization of mixed near-field and far-field sources

ICASSP 2017accepted

In this paper, 3D sound source localization method for simultaneously estimating both direction-of-arrival (DOA) and distance from the microphone array is proposed. For estimating distance, the off-grid problem must be overcome because the range of distance to be considered is quite broad and even n…

Cited by 0SourceScholar
2016

Physical-model based efficient data representation for many-channel microphone array

ICASSP 2016accepted

Recent development of microphone arrays which consist of more than several tens or hundreds microphones enables acquisition of rich spatial information of sound. Although such information possibly improve performance of any array signal processing technique, the amount of data will increase as the n…

Cited by 0SourceScholar
2015

Optically visualized sound field reconstruction based on sparse selection of point sound sources

ICASSP 2015accepted

Visualization is an effective way to understand the behavior of a sound field. There are several methods for such observation including optical measurement technique which enables a non-destructive acoustical observation by detecting density variation of the medium. For audible sound propagating thr…

Cited by 0SourceScholar
2015

Visualization of sound field by means of Schlieren method with spatio-temporal filtering

ICASSP 2015accepted

Visualization of sound field using Schlieren technique provides many advantages. It enables us to investigate the change of the sound field in real-time from every point of the observing region. However, since the density gradient of air caused by the disturbance of acoustic field is very small, it…

Cited by 0SourceScholar