ICASSP 2025accepted0 citations

A Domain-Specific Multilingual Speech Translation Corpus via Simultaneous Interpretation

Seunghee Han, Gary Geunbae Lee, Hung Soon Kim, Sunhee Kim, Minhwa Chung

Abstract

This paper presents a novel multilingual speech translation corpus for complex, domain-specific content in Korean, English, Spanish, and Japanese. The corpus contains 4,000 hours of parallel speech, including 1,000 hours of Korean audio with simultaneous sight interpretations in the other three languages by 294 professionals (242 interpreters and 52 Korean voice actors). It also includes transcriptions, translations, and annotations for all languages. The Dewey Decimal Classification was adapted to balance knowledge representation, and speech tasks were conducted in a controlled studio environment to ensure data consistency. Translation, transcription, and annotation workflows were managed through a custom-built platform. The corpus captures nuanced contexts, cultural sensitivities, and domain-specific terminology, addressing linguistic challenges like structural differences between SOV (Korean, Japanese) and SVO languages (English, Spanish). Preliminary evaluations indicate its potential to enhance end-to-end speech translation models, support cross-lingual transfer learning, and tackle real-time translation issues.

BibTeX
@inproceedings{icassp2025_adomainspecificm,
  title = {A Domain-Specific Multilingual Speech Translation Corpus via Simultaneous Interpretation},
  author = {Seunghee Han and Gary Geunbae Lee and Hung Soon Kim and Sunhee Kim and Minhwa Chung},
  booktitle = {ICASSP 2025},
  year = {2025}
}