COLING 2024main0 citations

FFSTC: Fongbe to French Speech Translation Corpus

D. Fortuné Kponou, Fréjus A. A. Laleye, Eugène Cokou Ezin

Abstract

In this paper, we introduce the Fongbe to French Speech Translation Corpus (FFSTC). This corpus encompasses approximately 31 hours of collected Fongbe language content, featuring both French transcriptions and corresponding Fongbe voice recordings. FFSTC represents a comprehensive dataset compiled through various collection methods and the efforts of dedicated individuals. Furthermore, we conduct baseline experiments using Fairseq’s transformer_s and conformer models to evaluate data quality and validity. Our results indicate a score BLEU of 8.96 for the transformer_s model and 8.14 for the conformer model, establishing a baseline for the FFSTC corpus.

BibTeX
@inproceedings{kponou-etal-2024-ffstc,
    title = "{FFSTC}: Fongbe to {F}rench Speech Translation Corpus",
    author = "Kponou, D. Fortun{\'e}  and
      Laleye, Fr{\'e}jus A. A.  and
      Ezin, Eug{\`e}ne Cokou",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.638/",
    pages = "7270--7276"
}