← Search

Akihiko Takashima

5 accepted papers

2021

Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss

ICASSP 2021accepted

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual speech separation is a technique to estimate the individual speech signals from a mi…

Cited by 0SourceScholar
2021

Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge Distillation

ICASSP 2021accepted

We present a novel large-context end-to-end automatic speech recognition (E2E-ASR) model and its effective training method based on knowledge distillation. Common E2E-ASR models have mainly focused on utterance-level processing in which each utterance is independently transcribed. On the other hand,…

Cited by 0SourceScholar
2021

MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-Training

ICASSP 2021accepted

This paper presents a self-supervised learning method for pointer-generator networks to improve spoken-text normalization. Spoken-text normalization that converts spoken-style text into style normalized text is becoming an important technology for improving subsequent processing such as machine tran…

Cited by 0SourceScholar
2020

Large-Context Pointer-Generator Networks for Spoken-to-Written Style Conversion

ICASSP 2020accepted

This paper introduces a spoken-to-written style conversion method that is suitable for handling a series of text such as discourses and conversations. Spoken-to-written style conversion can increase the readability of automatic speech recognition (ASR) outputs because ASR systems transcribe input sp…

Cited by 0SourceScholar
2020

Sequence-Level Consistency Training for Semi-Supervised End-to-End Automatic Speech Recognition

ICASSP 2020accepted

This paper presents a novel semi-supervised end-to-end automatic speech recognition (ASR) method that employs consistency training with the use of unlabeled data. In consistency training, unlabeled data can be utilized for constraining a model such that it becomes invariant to small deformation. In…

Cited by 0SourceScholar