← Search

Kazuya Takeda

12 accepted papers

2023

Learning to Predict Navigational Patterns From Partial Observations

RA-L 2023

Human beings cooperatively navigate rule-constrained environments by adhering to mutually known navigational patterns, which may be represented as directional pathways or road lanes. Inferring these navigational patterns from incompletely observed environments is required for intelligent mobile robo

Cited by 5SourcecodeScholar
2020

End-to-End Automatic Speech Recognition Integrated with CTC-Based Voice Activity Detection

ICASSP 2020accepted

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist temporal classification (CTC) and its extension of CTC/attention architectures. As opp…

Cited by 0SourceScholar
2020

Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

ICASSP 2020accepted

This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of- the-art E2E-TTS models, including Tacotron 2, Transformer TTS, and FastSpeech, and also provides recipes…

Cited by 0SourceScholar
2020

Weakly-Supervised Sound Event Detection with Self-Attention

ICASSP 2020accepted

In this paper, we propose a novel sound event detection (SED) method that incorporates a self-attention mechanism of the Transformer for a weakly-supervised learning scenario. The proposed method utilizes the Transformer encoder, which consists of multiple self-attention modules, allowing to take bo…

Cited by 0SourceScholar
2019

A Predictive Reward Function for Human-Like Driving Based on a Transition Model of Surrounding Environment

ICRA 2019poster

Driving is a complex task that requires the perception of the surrounding environment, decision making and control of the vehicle. Human drivers predict how surrounding objects move and decide an appropriate driving behavior. As with human drivers, autonomous driving vehicles should consider the con…

Cited by 5SourceScholar
2019

Point Cloud Compression for 3D LiDAR Sensor using Recurrent Neural Network with Residual Blocks

ICRA 2019poster

The use of 3D LiDAR, which has proven its capabilities in autonomous driving systems, is now expanding into many other fields. The sharing and transmission of point cloud data from 3D LiDAR sensors has broad application prospects in robotics. However, due to the sparseness and disorderly nature of t…

Cited by 111SourceScholar
2019

Scene-dependent Anomalous Acoustic-event Detection Based on Conditional Wavenet and I-vector

ICASSP 2019accepted

This paper proposes a scene-dependent anomalous acoustic-event detection based on conditional WaveNet and i-vector. The WaveNet builds normal acoustic event models by exhaustive learning of time-domain signals in the public space to provide scene-independent anomaly detection. I-vectors are used as…

Cited by 0SourceScholar
2017

BLSTM-HMM hybrid system combined with sound activity detection network for polyphonic Sound Event Detection

ICASSP 2017accepted

This paper presents a new hybrid approach for polyphonic Sound Event Detection (SED) which incorporates a temporal structure modeling technique based on a hidden Markov model (HMM) with a frame-by-frame detection method based on a bidirectional long short-term memory (BLSTM) recurrent neural network…

Cited by 0SourceScholar
2017

Music staging AI

ICASSP 2017accepted

Through smartphones, user enables to download/listen music anytime and anywhere. As a concept of a future audio player, we propose a framework of "music staging artificial intelligence (AI)". In that framework, audio object signals, e.g. vocal, guitar, bass, drums and keyboards, are assumed to be ex…

Cited by 0SourceScholar
2015

Exploring multi-channel features for denoising-autoencoder-based speech enhancement

ICASSP 2015accepted

This paper investigates a multi-channel denoising autoencoder (DAE)-based speech enhancement approach. In recent years, deep neural network (DNN)-based monaural speech enhancement and robust automatic speech recognition (ASR) approaches have attracted much attention due to their high performance. Al…

Cited by 109SourceScholar