← Search

Keisuke Imoto

15 accepted papers

2025

Formula-Supervised Sound Event Detection: Pre-Training Without Real Data

ICASSP 2025accepted

In this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effec…

Cited by 0SourceScholar
2025

Trainingless Adaptation of Pretrained Models for Environmental Sound Classification

ICASSP 2025accepted

Deep neural network (DNN)-based models for environmental sound classification are not robust against a domain to which training data do not belong, that is, out-of-distribution or unseen data. To utilize pretrained models for the unseen domain, adaptation methods, such as finetuning and transfer lea…

Cited by 0SourceScholar
2024

Environmental Sound Synthesis from Vocal Imitations and Sound Event Labels

ICASSP 2024accepted

One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as rhythm and pitch, using vocal imitations, which cannot be ex…

Cited by 0SourceScholar
2024

F1-EV score: Measuring The Likelihood of Estimating a Good Decision Threshold for Semi-Supervised Anomaly Detection

ICASSP 2024accepted

Anomalous sound detection (ASD) systems are usually compared by using threshold-independent performance measures such as AUCROC. However, for practical applications a decision threshold is needed to decide whether a given test sample is normal or anomalous. Estimating such a threshold is highly non-…

Cited by 0SourceScholar
2023

Visual Onoma-to-Wave: Environmental Sound Synthesis from Visual Onomatopoeias and Sound-Source Images

ICASSP 2023accepted

We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective, onoma-to-wave has been proposed to synthesize environmental s…

Cited by 0SourceScholar
2022

Environmental Sound Extraction Using Onomatopoeic Words

ICASSP 2022accepted

An onomatopoeic word, which is a character sequence that phonetically imitates a sound, is effective in expressing characteristics of sound such as duration, pitch, and timbre. We propose an environmental-sound-extraction method using onomatopoeic words to specify the target sound to be extracted. B…

Cited by 0SourceScholar
2022

Sound Event Detection Guided by Semantic Contexts of Scenes

ICASSP 2022accepted

Some studies have revealed that contexts of scenes (e.g., "home," "office," and "cooking") are advantageous for sound event detection (SED). Mobile devices and sensing technologies give useful information on scenes for SED without the use of acoustic signals. However, conventional methods can employ…

Cited by 0SourceScholar
2021

Impact of Sound Duration and Inactive Frames on Sound Event Detection Performance

ICASSP 2021accepted

In many methods of sound event detection (SED), a segmented time frame is regarded as one data sample to model training. The durations of sound events greatly depend on the sound event class, e.g., the sound event "fan" has a long duration, whereas the sound event "mouse clicking" is instantaneous.…

Cited by 0SourceScholar
2021

Sound Event Detection Based on Curriculum Learning Considering Learning Difficulty of Events

ICASSP 2021accepted

In conventional sound event detection (SED) models, two types of events, namely, those that are present and those that do not occur in an acoustic scene, are regarded as the same type of the events. The conventional SED methods cannot effectively exploit the difference between the two types of event…

Cited by 0SourceScholar
2020

Scene-Dependent Acoustic Event Detection with Scene Conditioning and Fake-Scene-Conditioned Loss

ICASSP 2020accepted

In this paper, we propose scene-dependent acoustic event detection (AED) with scene conditioning and fake-scene-conditioned loss. The proposed method employs a multitask network, that has not only AED part but also acoustic scene classification (ASC). The scenes predicted by ASC are employed as an a…

Cited by 0SourceScholar
2020

Sound Event Detection by Multitask Learning of Sound Events and Scenes with Soft Scene Labels

ICASSP 2020accepted

Sound event detection (SED) and acoustic scene classification (ASC) are major tasks in environmental sound analysis. Considering that sound events and scenes are closely related to each other, some works have addressed joint analyses of sound events and acoustic scenes based on multitask learning (M…

Cited by 0SourceScholar
2020

Sound Event Localization Based on Sound Intensity Vector Refined by Dnn-Based Denoising and Source Separation

ICASSP 2020accepted

We propose a direction-of-arrival (DOA) estimation method for Sound Event Localization and Detection (SELD). Direct estimation of DOA using a deep neural network (DNN), i.e. completely-datadriven approach, achieves high accuracy. However, there is a gap in the accuracy between DOA estimation for sin…

Cited by 0SourceScholar
2019

Joint Acoustic and Class Inference for Weakly Supervised Sound Event Detection

ICASSP 2019accepted

Sound event detection is a challenging task, especially for scenes with multiple simultaneous events. While event classification methods tend to be fairly accurate, event localization presents additional challenges, especially when large amounts of labeled data are not available. Task4 of the 2018 D…

Cited by 0SourceScholar
2019

Sound Event Detection Using Graph Laplacian Regularization Based on Event Co-occurrence

ICASSP 2019accepted

The types of sound events that occur in a situation are limited, and some sound events are likely to co-occur; for instance, "dishes" and "glass jingling." In this paper, we propose a technique of sound event detection utilizing graph Laplacian regularization taking the sound event co-occurrence int…

Cited by 0SourceScholar