← Search

Noboru Harada

30 accepted papers

2026

Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

ICASSP 2026poster

Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains…

Cited by 0SourcePDFScholar
2026

FedPM: Federated Learning Using Second-order Optimization with Preconditioned Mixing of Local Parameters

AAAI 2026technical

We propose Federated Preconditioned Mixing (FedPM), a novel Federated Learning (FL) method that leverages second-order optimization. Prior methods - such as LocalNewton, LTDA, and FedSophia - have incorporated second-order optimization in FL by performing iterative local updates on clients and apply

Cited by 0SourcePDFScholar
2025

Collision-less and Balanced Sampling for Language-Queried Audio Source Separation

ICASSP 2025accepted

Language-queried audio source separation (LASS) is an emerging research field that has recently received increasing attention. This task aims to isolate individual sources from a mixture of signals using natural language descriptions, enabling applications in various areas such as automatic audio ed…

Cited by 0SourceScholar
2025

Sound Source Distance Estimation Utilizing Physics-informed Prior for Sound Event Localization and Detection

ICASSP 2025accepted

Sound Event Localization and Detection (SELD) is the combined task of detecting sound events and estimating their spatial locations. We propose a Sound source Distance Estimation (SDE) method for SELD that utilizes a physics-informed prior. The conventional data-driven approach of SDE for SELD can h…

Cited by 0SourceScholar
2025

Spatial Annotation-free Training for Sound Event Localization and Detection

ICASSP 2025accepted

Sound Event Localization and Detection (SELD) is the task of estimating the class, duration, and direction of arrival (DOA) of sound events. State-of-the-art SELD systems use a data-driven approach based on Deep Neural Networks (DNNs) to deal with complex situations involving overlapping and moving…

Cited by 1SourceScholar
2025

Stereo Downmix in 3GPP IVAS for EVS Compatibility

ICASSP 2025accepted

The 3GPP IVAS codec specifies an EVS-compatible stereo downmix as one of the key functionalities. This paper describes how this novel active downmix scheme has been devised to achieve high and stable quality from stereo input to EVS encoder/decoder with no additional algorithmic delay. An example of…

Cited by 0SourceScholar
2024

6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on Self-Motioning Human

ICASSP 2024accepted

We aim to perform sound event localization and detection (SELD) using wearable equipment for a moving human, such as a pedestrian. Conventional SELD tasks have dealt only with microphone arrays located in static positions. However, self-motion with three rotational and three translational degrees of…

Cited by 0SourceScholar
2024

Unrestricted Global Phase Bias-Aware Single-Channel Speech Enhancement with Conformer-Based Metric Gan

ICASSP 2024accepted

With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally outstanding. However, enhancing the phase spectrum using neural networks is often…

Cited by 0SourceScholar
2023

Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input

ICASSP 2023accepted

Masked Autoencoders is a simple yet powerful self-supervised learning method. However, it learns representations indirectly by reconstructing masked input patches. Several methods learn representations directly by predicting representations of masked patches; however, we think using all patches to e…

Cited by 0SourceScholar
2022

Multi-View And Multi-Modal Event Detection Utilizing Transformer-Based Multi-Sensor Fusion

ICASSP 2022accepted

We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microphones and their weak labels. In this task, distributed sensors are utilized complementarily to capture events that are di…

Cited by 0SourceScholar
2021

Asynchronous Decentralized Optimization With Implicit Stochastic Variance Reduction

ICML 2021spotlight

A novel asynchronous decentralized optimization method that follows Stochastic Variance Reduction (SVR) is proposed. Average consensus algorithms, such as Decentralized Stochastic Gradient Descent (DSGD), facilitate distributed training of machine learning models. However, the gradient will drift wi…

2020

A Frequency-Domain BSS Method Based on ℓ1 Norm, Unitary Constraint, and Cayley Transform

ICASSP 2020accepted

We propose a frequency-domain blind source separation method that uses (a) the ℓ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> norm of orthonormal vectors of estimated source signals as a sparsity measure and (b) Cayley transform for optimizin…

Cited by 0SourceScholar
2020

Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement

ICASSP 2020accepted

We propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-t…

Cited by 0SourceScholar
2020

Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks

ICASSP 2020accepted

Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data, several studies presented deep neural network (DNN)–based phase…

Cited by 0SourceScholar
2020

Real-Time Speech Enhancement Using Equilibriated RNN

ICASSP 2020accepted

We propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capabili…

Cited by 44SourceScholar
2020

SPIDERnet: Attention Network For One-Shot Anomaly Detection In Sounds

ICASSP 2020accepted

We propose a similarity function for one-shot anomaly detection in sounds (ADS) called SPecific anomaly IDentifiER network (SPIDERnet). In ADS systems, since overlooking an anomaly may result in serious incidents, we need to update such systems using an (often only one) overlooked anomalous sample.…

Cited by 0SourceScholar
2020

Subjective Quality Estimation Using PESQ For Hands-Free Terminals

ICASSP 2020accepted

Previous reports have mentioned the possibility that subjective quality of the echo-suppressed speech signal can be estimated based on perceptual evaluation of speech quality (PESQ), but there are few experimental results. We propose third-party listening and conversational test procedures to assess…

Cited by 0SourceScholar
2019

A Two-class Hyper-spherical Autoencoder for Supervised Anomaly Detection

ICASSP 2019accepted

Supervised anomaly detection has been a tough problem due to its necessity of special handling of unseen anomalies. In this paper, we present a heuristic implementation of variational auto-encoder with von-Mises Fisher prior applied to a supervised anomaly detector. The closed latent space like sphe…

Cited by 0SourceScholar
2019

AdaFlow: Domain-adaptive Density Estimator with Application to Anomaly Detection and Unpaired Cross-domain Translation

ICASSP 2019accepted

We tackle unsupervised anomaly detection (UAD), a problem of detecting data that significantly differ from normal data. UAD is typically solved by using density estimation. Recently, deep neural network (DNN)-based density estimators, such as Normalizing Flows, have been attracting attention. Howeve…

Cited by 0SourceScholar
2019

Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement

ICASSP 2019accepted

We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when…

Cited by 0SourceScholar
2019

Function Designable Beamformer Based on Probabilistic Assumptions on Filter and Its Auxiliary Variables

ICASSP 2019accepted

We propose a novel beamformer design method that exploits probabilistic assumptions on auxiliary variables derived from filters and observed signals. Many conventional beamformer design methods can be understood in the context of optimization problems for some probabilistic cost functions. However,…

Cited by 0SourceScholar
2019

SNIPER: Few-shot Learning for Anomaly Detection to Minimize False-negative Rate with Ensured True-positive Rate

ICASSP 2019accepted

In anomaly detection systems, overlooking anomalies may result in serious incidents. Thus, when a system overlooks an anomaly, we need to update the system to never overlook the observed type of anomalies twice. There are roughly two possible approaches to solve this problem; re-training the whole s…

Cited by 0SourceScholar
2018

Complementary Set Variational Autoencoder for Supervised Anomaly Detection

ICASSP 2018accepted

Anomalies have broad patterns corresponding to their causes. In industry, anomalies are typically observed as equipment failures. Anomaly detection aims to detect such failures as anomalies. Although this is usually a binary classification task, the potential existence of unseen (unknown) failures m…

Cited by 0SourceScholar
2018

End-to-End Sound Source Enhancement Using Deep Neural Network in the Modified Discrete Cosine Transform Domain

ICASSP 2018accepted

This paper presents an end-to-end deep neural network (DNN)-based source enhancement on the basis of a time-frequency (T-F) mask processing in the modified discrete cosine transform (MDCT)-domain. To retrieve the target signal perfectly in the discrete Fourier transform (DFT)-domain, both amplitude…

Cited by 0SourceScholar
2015

Standardization of the new 3GPP EVS codec

ICASSP 2015accepted

A new codec for Enhanced Voice Services (EVS), the successor of the current mobile HD voice codec AMR-WB, was standardized by the 3rd Generation Partnership Project (3GPP) in September 2014. The EVS codec addresses 3GPP's needs for cutting-edge technology enabling operation of 3GPP mobile communicat…

Cited by 0SourceScholar