← Search

Motoi Omachi

5 accepted papers

2023

Align, Write, Re-Order: Explainable End-to-End Speech Translation via Operation Sequence Generation

ICASSP 2023accepted

The black-box nature of end-to-end speech-to-text translation (E2E ST) makes it difficult to understand how source language inputs are being mapped to the target language. To solve this problem, we propose to simultaneously generate automatic speech recognition (ASR) and ST predictions such that eac…

Cited by 0SourceScholar
2022

Non-Autoregressive End-To-End Automatic Speech Recognition Incorporating Downstream Natural Language Processing

ICASSP 2022accepted

We propose a fast and accurate end-to-end (E2E) model, which executes automatic speech recognition (ASR) and downstream natural language processing (NLP) simultaneously. The proposed approach predicts a single-aligned sequence of transcriptions and linguistic annotations such as part-of-speech (POS)…

Cited by 0SourceScholar
2021

End-to-end ASR to jointly predict transcriptions and linguistic annotations

NAACL 2021long

We propose a Transformer-based sequence-to-sequence model for automatic speech recognition (ASR) capable of simultaneously transcribing and annotating audio with linguistic information such as phonemic transcripts or part-of-speech (POS) tags. Since linguistic information is important in natural lan…

Cited by 13SourcePDFScholar
2020

Attention-Based ASR with Lightweight and Dynamic Convolutions

ICASSP 2020accepted

End-to-end (E2E) automatic speech recognition (ASR) with sequence-to-sequence models has gained attention because of its simple model training compared with conventional hidden Markov model based ASR. Recently, several studies report the state-of-the-art E2E ASR results obtained by Transformer. Comp…

Cited by 0SourceScholar
2018

Multi Scale Feedback Connection for Noise Robust Acoustic Modeling

ICASSP 2018accepted

Simply feeding of a last hidden layer of the deep neural network (DNN) back to the input layer recently found to be effective for noise robust acoustic modeling. Such high level feature strengthens the robustness of DNN based acoustic model while paying approximately twice the computational cost. In…

Cited by 0SourceScholar