2025
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
Andrei-Marius Avram, B{\u{a}}nescu Ema-Ioana, Anda-Teodora Robea, Dumitru-Clementin Cercel, Mihaela-Claudia Cercel
EMNLP 2025
This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced between the Romanian language spoken in Romania and the Republic of Moldova. We further propose a multi-target adversarial