← Search

Kazuki Shimada

11 accepted papers

2026

SAVGBENCH: BENCHMARKING SPATIALLY ALIGNED AUDIO-VIDEO GENERATION

ICASSP 2026poster

This work addresses the lack of multimodal generative models capable of producing high-quality videos with spatially aligned audio. While recent advancements in generative models have been successful in video generation, they often overlook the spatial alignment between audio and visuals, which is e…

Cited by 0SourcePDFScholar
2024

Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

ICASSP 2024accepted

Diffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusion process with enhanced features estimated by a predictive SE system. However, the pipeline structure currently does not…

Cited by 0SourceScholar
2024

Zero- and Few-Shot Sound Event Localization and Detection

ICASSP 2024accepted

Sound event localization and detection (SELD) systems estimate direction-of-arrival (DOA) and temporal activation for sets of target classes. Neural network (NN)-based SELD systems have performed well in various sets of target classes, but they only output the DOA and temporal activation of preset c…

Cited by 0SourceScholar
2023

An Attention-Based Approach to Hierarchical Multi-Label Music Instrument Classification

ICASSP 2023accepted

Although music is typically multi-label, many works have studied hierarchical music tagging with simplified settings such as single-label data. Moreover, there lacks a framework to describe various joint training methods under the multi-label setting. In order to discuss the above topics, we introdu…

Cited by 0SourceScholar
2023

STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events

NeurIPS 2023poster

While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from the feet of a walker. This paper proposes an audio-visual sou…

2022

Multi-ACCDOA: Localizing And Detecting Overlapping Sounds From The Same Class With Auxiliary Duplicating Permutation Invariant Training

ICASSP 2022accepted

Sound event localization and detection (SELD) involves identifying the direction-of-arrival (DOA) and the event class. The SELD methods with a class-wise output format make the model predict activities of all sound event classes and corresponding locations. The class-wise methods can output activity…

Cited by 111SourceScholar
2022

Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection

ICASSP 2022accepted

Recording and annotating real sound events for a sound event localization and detection (SELD) task is time consuming, and data augmentation techniques are often favored when the amount of data is limited. However, how to augment the spatial information in a dataset, including unlabeled directional…

Cited by 0SourceScholar
2022

Spatial Mixup: Directional Loudness Modification as Data Augmentation for Sound Event Localization and Detection

ICASSP 2022accepted

Data augmentation methods have shown great importance in diverse supervised learning problems where labeled data is scarce or costly to obtain. For sound event localization and detection (SELD) tasks several augmentation methods have been proposed, with most borrowing ideas from other domains such a…

Cited by 0SourceScholar
2021

Accdoa: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization And Detection

ICASSP 2021accepted

Neural-network (NN)-based methods show high performance in sound event localization and detection (SELD). Conventional NN-based methods use two branches for a sound event detection (SED) target and a direction-of-arrival (DOA) target. The two-branch representation with a single network has to decide…

Cited by 0SourceScholar
2020

Metric Learning with Background Noise Class for Few-Shot Detection of Rare Sound Events

ICASSP 2020accepted

Few-shot learning systems for sound event recognition have gained interests since they require only a few examples to adapt to new target classes without fine-tuning. However, such systems have only been applied to chunks of sounds for classification or verification. In this paper, we aim to achieve…

Cited by 0SourceScholar
2018

Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech Recognition

ICASSP 2018accepted

This paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to…

Cited by 0SourceScholar