← Search

Yushi Aono

5 accepted papers

2022

Bilateral Video Magnification Filter

CVPR 2022poster

Eulerian video magnification (EVM) has progressed to magnify subtle motions with a target frequency even under the presence of large motions of objects. However, existing EVM methods often fail to produce desirable results in real videos due to (1) mis-extracting subtle motions with a non-target fre…

Cited by 11PDFScholar
2019

Large Context End-to-end Automatic Speech Recognition via Extension of Hierarchical Recurrent Encoder-decoder Models

ICASSP 2019accepted

This paper describes a novel end-to-end automatic speech recognition (ASR) method that takes into consideration long-range sequential context information beyond utterance boundaries. In spontaneous ASR tasks such as those for discourses and conversations, the input speech often comprises a series of…

Cited by 0SourceScholar
2018

Soft-Target Training with Ambiguous Emotional Utterances for DNN-Based Speech Emotion Classification

ICASSP 2018accepted

This paper presents a novel emotion classification method for natural speech. One of the problems in the state-of-the-art method based on Deep Neural Network (DNN) is the paucity of the training data compared to model complexity. To solve this problem, this paper utilizes the ambiguous emotional utt…

Cited by 0SourceScholar
2017

Domain adaptation of DNN acoustic models using knowledge distillation

ICASSP 2017accepted

Constructing deep neural network (DNN) acoustic models from limited training data is an important issue for the development of automatic speech recognition (ASR) applications that will be used in various application-specific acoustic environments. To this end, domain adaptation techniques that train…

Cited by 0SourceScholar
2017

Parallel phonetically aware DNNs and LSTM-RNNS for frame-by-frame discriminative modeling of spoken language identification

ICASSP 2017accepted

Parallel phonetically aware deep neural networks (PPA-DNNs) and long short-term memory recurrent neural networks (PPA-LSTM-RNNs) to enhance frame-by-frame discriminative modeling of spoken language identification are proposed. This idea is inspired by traditional systems based on parallel phoneme re…

Cited by 0SourceScholar