← Search

Shuichiro Shimizu

5 accepted papers

2025

ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems

NAACL 2025system demonstrations

Advancements in audio foundation models (FMs) have fueled interest in end-to-end (E2E) spoken dialogue systems, but different web interfaces for each system makes it challenging to compare and contrast them effectively. Motivated by this, we introduce an open-source, user-friendly toolkit designed t…

2025

When Large Language Models Meet Speech: A Survey on Integration Approaches

ACL 2025finding

Recent advancements in large language models (LLMs) have spurred interest in expanding their application beyond text-based tasks. A large number of studies have explored integrating other modalities with LLMs, notably speech modality, which is naturally related to text. This paper surveys the integr…

Cited by 0SourcePDFScholar
2024

MELD-ST: An Emotion-aware Speech Translation Dataset

ACL 2024findings

Emotion plays a crucial role in human conversation. This paper underscores the significance of considering emotion in speech translation. We present the MELD-ST dataset for the emotion-aware speech translation task, comprising English-to-Japanese and English-to-German language pairs. Each language p…

Cited by 2SourcePDFScholar
2023

Towards Speech Dialogue Translation Mediating Speakers of Different Languages

ACL 2023findings

We present a new task, speech dialogue translation mediating speakers of different languages. We construct the SpeechBSD dataset for the task and conduct baseline experiments. Furthermore, we consider context to be an important aspect that needs to be addressed in this task and propose two ways of u…

2023

Video-Helpful Multimodal Machine Translation

EMNLP 2023long main

Existing multimodal machine translation (MMT) datasets consist of images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity, making visual information ineffective in generating appropriate translations. Recent work has constructed an ambiguous subtitles da…

Cited by 0SourcecodeScholar