← Search

Kazuhiro Nakadai

29 accepted papers

2026

Unsupervised Single-Channel Audio Separation with Diffusion Source Priors

AAAI 2026technical

Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated paired data. However, obtaining high-quality paired data in real-world scenarios is often difficult. This data scarcity ca

Cited by 0SourcePDFScholar
2025

Improvement in Sign Language Translation Using Text CTC Alignment

COLING 2025main

Current sign language translation (SLT) approaches often rely on gloss-based supervision with Connectionist Temporal Classification (CTC), limiting their ability to handle non-monotonic alignments between sign language video and spoken text. In this work, we propose a novel method combining joint CT…

2025

Multilingual Gloss-free Sign Language Translation: Towards Building a Sign Language Foundation Model

ACL 2025short

Sign Language Translation (SLT) aims to convert sign language (SL) videos into spoken language text, thereby bridging the communication gap between the sign and the spoken community. While most existing works focus on translating a single SL into a single spoken language (one-to-one SLT), leveraging…

2025

Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments

IROS 2025

Accurately estimating sound source positions is crucial for robot audition. However, existing sound source localization methods typically rely on a microphone array with at least two spatially preconfigured microphones. This requirement hinders the applicability of microphone-based robot audition sy

Cited by 0SourcecodeScholar
2025

Swarm Active Audition with Robots and Drones: Real-World Performance Validation

IROS 2025

Search and rescue (SAR) operations in large-scale disaster sites, such as areas affected by earthquakes, require rapid victim detection. While drones equipped with cameras are commonly used for SAR, their effectiveness is limited in visually obstructed environments, because of debris, smoke, or fog.

Cited by 0SourceScholar
2022

Outdoor evaluation of sound source localization for drone groups using microphone arrays

IROS 2022poster

For robot and drone auditions, microphone arrays have been used for estimating sound source directions and sound source locations. By using sound source localization techniques, for example, drones can detect people calling for help even if the target person is not visible. Most sound source localiz…

Cited by 3SourceScholar
2022

Spotforming by NMF Using Multiple Microphone Arrays

IROS 2022poster

Sound source separation is a method to extract a target sound source from a mixture of various sound sources and noises. One of the typical sound source separation methods is beamforming, which can separate sound sources by direction based on the phase difference between channels from the recorded s…

Cited by 1SourceScholar
2021

Fully-Online Always-Adaptation of Transfer Functions and Its Application to Sound Source Localization and Separation

IROS 2021poster

This paper addresses fully-online always-adaptation of a transfer function for robot audition systems based on microphone array processing. The transfer function represents signal propagation characteristics between a microphone and a sound source, which provides essential information for real-world…

Cited by 4SourceScholar
2020

Synchronization of Microphones Based on Rank Minimization of Warped Spectrum for Asynchronous Distributed Recording

IROS 2020poster

This paper describes a new method for synchronizing microphones based on spectral warping in an asynchronous microphone array. In an audio signal observed by an asynchronous microphone array, two factors are involved: the time lag caused by a mismatch of the sampling rate and offset between micropho…

Cited by 5SourceScholar
2019

An Integrated Framework for Field Recording, Localization, Classification and Annotation of Birdsongs Using Robot Audition Techniques - Harkbird 2.0

ICASSP 2019accepted

Bird vocalizations are one of the important subjects in ecoacoustics because birds communicate diversely using various vocalizations such as songs and calls. We have developed a portable system, HARKBird to provide a basic function, i.e., birdsong localization, which automatically extracts sound sou…

Cited by 0SourceScholar
2018

Extracting the Relationship between the Spatial Distribution and Types of Bird Vocalizations Using Robot Audition System HARK

IROS 2018poster

For a deeper understanding of ecological functions and semantics of wild bird vocalizations (i.e., songs and calls), it is important to clarify the fine-scaled and detailed relationships among their characteristics of vocalizations and their behavioral contexts. However, it takes a lot of time and e…

Cited by 13SourceScholar
2018

HARK-Bird-Box: A Portable Real-time Bird Song Scene Analysis System

IROS 2018poster

This paper addresses real-time bird song scene analysis. Observation of animal behavior such as communication of wild birds would be aided by a portable device implementing a real-time system that can localize sound sources, measure their timing, classify their sources, and visualize these factors o…

Cited by 14SourceScholar
2018

Multi-timescale Feature-extraction Architecture of Deep Neural Networks for Acoustic Model Training from Raw Speech Signal

IROS 2018poster

This paper describes a new architecture of deep neural networks (DNNs) for acoustic models. Training DNNs from raw speech signals will provide 1) novel features of signals, 2) normalization-free processing such as utterance-wise mean subtraction, and 3) low-latency speech recognition for robot audit…

Cited by 7SourceScholar
2017

Development of microphone-array-embedded UAV for search and rescue task

IROS 2017poster

This paper addresses online outdoor sound source localization using a microphone array embedded in an unmanned aerial vehicle (UAV). In addition to sound source localization, sound source enhancement and robust communication method are also described. This system is one instance of deployment of our…

Cited by 57SourceScholar
2016

Online simultaneous localization and mapping of multiple sound sources and asynchronous microphone arrays

IROS 2016poster

This paper presents an online method of simultaneous localization and mapping (SLAM) for estimating the positions of multiple moving sound sources and stationary robots and synchronizing microphone arrays attached to those robots. Since each robot with a microphone array can solely estimate the dire…

Cited by 18SourceScholar
2016

Partially Shared Deep Neural Network in sound source separation and identification using a UAV-embedded microphone array

IROS 2016poster

This paper addresses sound source separation and identification for noise-contaminated acoustic signals recorded with a microphone array embedded in an Unmanned Aerial Vehicle (UAV), aiming at people's voice detection quickly and widely in a disaster situation. The key approach to achieve this is De…

Cited by 49SourceScholar
2016

Robust sound source mapping using three-layered selective audio rays for mobile robots

IROS 2016poster

This paper investigates sound source mapping in a real environment using a mobile robot. Our approach is based on audio ray tracing which integrates occupancy grids and sound source localization using a laser range finder and a microphone array. Previous audio ray tracing approaches rely on all obse…

Cited by 11SourceScholar
2016

Semi-automatic bird song analysis by spatial-cue-based integration of sound source detection, localization, separation, and identification

IROS 2016poster

This paper addresses bird song analysis based on semi-automatic annotation. Research in animal behavior, especially with birds, would be aided by automated (or semiautomated) systems that can localize sounds, measure their timing, and identify their source. This is difficult to achieve in real envir…

Cited by 26SourceScholar
2015

Audio-visual scene understanding utilizing text information for a cooking support robot

IROS 2015poster

This paper addresses multimodal “scene understanding” for a robot using audio-visual and text information. Scene understanding is defined by extracting six-W information such as What, When, Where, Who, Why, and hoW on the surrounding environment. Although scene understanding for a robot has been stu…

Cited by 15SourceScholar
2015

Interactive sound source localization using robot audition for tablet devices

IROS 2015poster

This paper investigates localization of sound sources in a real environment using a tablet device. For the localization, we use build-in sensors on a tablet device and additionally mount a cover with a microphone array. Because of the flat shape and limited sensor performance, the localization has m…

Cited by 5SourceScholar
2015

Microphone-accelerometer based 3D posture estimation for a hose-shaped rescue robot

IROS 2015poster

3D posture estimation for a hose-shaped robot is critical in rescue activities due to complex physical environments. Conventional sound-based posture estimation assumes rather flat physical environments and focuses only on 2D, resulting in poor performance in real world environments with rubble. Thi…

Cited by 16SourceScholar
2015

On-the-spot calibration of microphone array Transfer Functions for robot audition

ICRA 2015poster

This paper investigates the calibration of a microphone array based robot audition system, namely calibration of microphone array Transfer Functions (TFs). There are mainly two methods to obtain TFs: geometrical calculation and measurement. The geometrical calculation has difficulty in simulating ro…

Cited by 7SourceScholar
2015

Robot audition based Acoustic Event Identification using a Bayesian model considering spectral and temporal uncertainties

IROS 2015poster

To analyze auditory scenes of robots' surrounding environments, not only speeches but also non-speech sounds are important, which are spatially distributed and have different spectral and temporal characteristics. Thus, this paper investigates Acoustic Event Identification (AEI) which includes probl…

Cited by 9SourceScholar
2015

Temporal smearing compensation in reverberant environment for speech-based human-robot interaction

ICRA 2015poster

Speech-based human-robot interaction is often plagued with issues such as reverberation and changes in speaker position that impacts overall performance. In this paper, we show a method in compensating the joint effects of reverberation and the change in speaker position. The acoustic perturbation c…

Cited by 2SourceScholar
2015

Utilizing visual cues in robot audition for sound source discrimination in speech-based human-robot communication

IROS 2015poster

It is easy for human beings to discern whether an observed acoustic signal is a direct speech, reflected speech or noise through simple listening. Relying purely on acoustic cues is enough for human beings to discriminate between the different kinds of sound sources which is not straightforward for…

Cited by 6SourceScholar