← Search

Masahiro Yasuda

12 accepted papers

2025

Collision-less and Balanced Sampling for Language-Queried Audio Source Separation

ICASSP 2025accepted

Language-queried audio source separation (LASS) is an emerging research field that has recently received increasing attention. This task aims to isolate individual sources from a mixture of signals using natural language descriptions, enabling applications in various areas such as automatic audio ed…

Cited by 0SourceScholar
2025

Sound Source Distance Estimation Utilizing Physics-informed Prior for Sound Event Localization and Detection

ICASSP 2025accepted

Sound Event Localization and Detection (SELD) is the combined task of detecting sound events and estimating their spatial locations. We propose a Sound source Distance Estimation (SDE) method for SELD that utilizes a physics-informed prior. The conventional data-driven approach of SDE for SELD can h…

Cited by 0SourceScholar
2025

Spatial Annotation-free Training for Sound Event Localization and Detection

ICASSP 2025accepted

Sound Event Localization and Detection (SELD) is the task of estimating the class, duration, and direction of arrival (DOA) of sound events. State-of-the-art SELD systems use a data-driven approach based on Deep Neural Networks (DNNs) to deal with complex situations involving overlapping and moving…

Cited by 1SourceScholar
2024

6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on Self-Motioning Human

ICASSP 2024accepted

We aim to perform sound event localization and detection (SELD) using wearable equipment for a moving human, such as a pedestrian. Conventional SELD tasks have dealt only with microphone arrays located in static positions. However, self-motion with three rotational and three translational degrees of…

Cited by 0SourceScholar
2024

Online Target Sound Extraction with Knowledge Distillation from Partially Non-Causal Teacher

ICASSP 2024accepted

Target Sound Extraction (TSE) is a technique for extracting sound events belonging to a target sound class in a mixture using a Deep Neural Network (DNN). Offline TSE that uses non-causal models has achieved high extraction performance. However, many applications require online processing. Simply co…

Cited by 10SourceScholar
2022

APPLADE: Adjustable Plug-and-Play Audio Declipper Combining DNN with Sparse Optimization

ICASSP 2022accepted

In this paper, we propose an audio declipping method that takes advantages of both sparse optimization and deep learning. Since sparsity-based audio declipping methods have been developed upon constrained optimization, they are adjustable and well-studied in theory. However, they always uniformly pr…

Cited by 0SourceScholar
2022

Echo-Aware Adaptation of Sound Event Localization and Detection in Unknown Environments

ICASSP 2022accepted

Our goal is to develop a sound event localization and detection (SELD) system that works robustly in unknown environments. A SELD system trained on known environment data is degraded in an unknown environment due to environmental effects such as reverberation and noise not contained in the training…

Cited by 0SourceScholar
2022

Multi-View And Multi-Modal Event Detection Utilizing Transformer-Based Multi-Sensor Fusion

ICASSP 2022accepted

We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microphones and their weak labels. In this task, distributed sensors are utilized complementarily to capture events that are di…

Cited by 0SourceScholar
2022

Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around Head

ICASSP 2022accepted

Sound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with the sound signals observed by a microphone array. Although ambisonic microphones are popular in the literature of SELD, t…

Cited by 0SourceScholar
2020

SPIDERnet: Attention Network For One-Shot Anomaly Detection In Sounds

ICASSP 2020accepted

We propose a similarity function for one-shot anomaly detection in sounds (ADS) called SPecific anomaly IDentifiER network (SPIDERnet). In ADS systems, since overlooking an anomaly may result in serious incidents, we need to update such systems using an (often only one) overlooked anomalous sample.…

Cited by 0SourceScholar
2020

Sound Event Detection by Multitask Learning of Sound Events and Scenes with Soft Scene Labels

ICASSP 2020accepted

Sound event detection (SED) and acoustic scene classification (ASC) are major tasks in environmental sound analysis. Considering that sound events and scenes are closely related to each other, some works have addressed joint analyses of sound events and acoustic scenes based on multitask learning (M…

Cited by 0SourceScholar
2020

Sound Event Localization Based on Sound Intensity Vector Refined by Dnn-Based Denoising and Source Separation

ICASSP 2020accepted

We propose a direction-of-arrival (DOA) estimation method for Sound Event Localization and Detection (SELD). Direct estimation of DOA using a deep neural network (DNN), i.e. completely-datadriven approach, achieves high accuracy. However, there is a gap in the accuracy between DOA estimation for sin…

Cited by 0SourceScholar