← Search

Hirofumi Inaguma

17 accepted papers

2025

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

EMNLP 2025

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration appro

2024

Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction

ICLR 2024spotlight

Existing Self-Supervised Learning (SSL) models for speech typically process speech signals at a fixed resolution of 20 milliseconds. This approach overlooks the varying informational content present at different resolutions in speech signals. In contrast, this paper aims to incorporate multi-resolut…

Cited by 26SourcePDFScholar
2023

Enhancing Speech-To-Speech Translation with Multiple TTS Targets

ICASSP 2023accepted

It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both source and target speech. Therefore to train a direct S2ST system, previous works usually utilize text-to-speech (TTS) syst…

Cited by 0SourceScholar
2023

Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks

ACL 2023long

Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits and drawbacks for speech-to-text tasks. In order to leverage strengths of both modeling methods, we propose a solution…

2023

Named Entity Detection and Injection for Direct Speech Translation

ICASSP 2023accepted

In a sentence, certain words are critical for its semantic. Among them, named entities (NEs) are notoriously challenging for neural models. Despite their importance, their accurate handling has been neglected in speech-to-text (S2T) translation research, and recent work has shown that S2T models per…

Cited by 0SourceScholar
2023

Simple and Effective Unsupervised Speech Translation

ACL 2023long

The amount of labeled data to train models for speech tasks is limited for most languages, however, the data scarcity is exacerbated for speech translation which requires labeled data covering two different languages. To address this issue, we study a simple and effective approach to build speech tr…

2023

Speech-to-Speech Translation for a Real-world Unwritten Language

ACL 2023findings

We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard text writing systems. We use English-Taiwanese Hokkien as a case study, and present an end-to-end solution from training d…

2023

UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units

ACL 2023long

Direct speech-to-speech translation (S2ST), in which all components can be optimized jointly, is advantageous over cascaded approaches to achieve fast inference with a simplified pipeline. We present a novel two-pass direct S2ST architecture, UnitY, which first generates textual representations and…

2021

Improved Mask-CTC for Non-Autoregressive End-to-End ASR

ICASSP 2021accepted

For real-world deployment of automatic speech recognition (ASR), the system is desired to be capable of fast inference while relieving the requirement of computational resources. The recently proposed end-to-end ASR system based on mask-predict with connectionist temporal classification (CTC), Mask-…

Cited by 0SourceScholar
2021

ORTHROS: non-autoregressive end-to-end speech translation With dual-decoder

ICASSP 2021accepted

Fast inference speed is an important goal towards real-world deployment of speech translation (ST) systems. End-to-end (E2E) models based on the encoder-decoder architecture are more suitable for this goal than traditional cascaded systems, but their effectiveness regarding decoding speed has not be…

Cited by 0SourceScholar
2021

Recent Developments on Espnet Toolkit Boosted By Conformer

ICASSP 2021accepted

In this study, we present recent developments on ESPnet: End-to- End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a wide range of end- to-end speech processing applications, suc…

Cited by 0SourceScholar
2021

Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation

NAACL 2021long

A conventional approach to improving the performance of end-to-end speech translation (E2E-ST) models is to leverage the source transcription via pre-training and joint training with automatic speech recognition (ASR) and neural machine translation (NMT) tasks. However, since the input modalities ar…

2020

Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR

ICASSP 2020accepted

Recently, a few novel streaming attention-based sequence-to-sequence (S2S) models have been proposed to perform online speech recognition with linear-time decoding complexity. However, in these models, the decisions to generate tokens are delayed compared to the actual acoustic boundaries since thei…

Cited by 0SourceScholar
2019

Language Model Integration Based on Memory Control for Sequence to Sequence Speech Recognition

ICASSP 2019accepted

In this paper, we explore several new schemes to train a seq2seq model to integrate a pre-trained language model (LM). Our proposed fusion methods focus on the memory cell state and the hidden state in the seq2seq decoder long short-term memory (LSTM), and the memory cell state is updated by the LM…

Cited by 0SourceScholar
2019

Transfer Learning of Language-independent End-to-end ASR with Language Model Fusion

ICASSP 2019accepted

This work explores better adaptation methods to low-resource languages using an external language model (LM) under the framework of transfer learning. We first build a language-independent ASR system in a unified sequence-to-sequence (S2S) architecture with a shared vocabulary among all languages. D…

Cited by 0SourceScholar
2018

Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based Model

ICASSP 2018accepted

This paper addresses end-to-end speech recognition which directly maps acoustic features to a word sequence. The acoustic-to-word model is attractive since it does not require an external language model and an elaborate decoder, resulting in extremely simple and fast decoding. The apparent drawback…

Cited by 0SourceScholar
2018

An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition

ICASSP 2018accepted

Social signals such as laughter and fillers are often observed in natural conversation, and they play various roles in human-to-human communication. Detecting these events is useful for transcription systems to generate rich transcription and for dialogue systems to behave as we do such as synchroni…

Cited by 0SourceScholar