COLING 2024main0 citations

Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech

Martijn Bentum, Eric Sanders, Antal P.J. van den Bosch, Douwe Zeldenrust, Henk van den Heuvel

Abstract

The Dutch Dialect Database (also known as the ‘Nederlandse Dialectenbank’) contains dialectal variations of Dutch that were recorded all over the Netherlands in the second half of the twentieth century. A subset of these recordings of about 300 hours were enriched with manual orthographic transcriptions, using non-standard approximations of dialectal speech. In this paper we describe the creation of a corpus containing both the audio recordings and their corresponding transcriptions and focus on our method for aligning the recordings with the transcriptions and the metadata.

BibTeX
@inproceedings{bentum-etal-2024-corpus,
    title = "Corpus Creation and Automatic Alignment of Historical {D}utch Dialect Speech",
    author = "Bentum, Martijn  and
      Sanders, Eric  and
      van den Bosch, Antal P.J.  and
      Zeldenrust, Douwe  and
      van den Heuvel, Henk",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.357/",
    pages = "4021--4029"
}