SADA: Saudi Audio Dataset for Arabic
Sadeen Alharbi, Areeb Alowisheq, Zoltán Tüske, Kareem Darwish, Abdullah Alrajeh, Abdulmajeed Alrowithi, Aljawharah Bin Tamran, Asma Ibrahim
Abstract
Arabic is among the most challenging languages in the world. Unfortunately, the scarcity of Arabic datasets makes studies in Arabic speech technology demanding. This paper introduces SADA, the Saudi Audio Dataset for Arabic, with 668 hours of high-quality audio suitable for supervised training. The audio recordings were sourced from 57 television shows provided by the Saudi Broadcasting Authority. The audio covers both read and spontaneous speaking styles in various genres. The National Center for Artificial Intelligence in Saudi Arabia transcribed and prepared the data for training and processing. The recordings are in Arabic. Most are in Saudi dialects, while other Arabic dialects include Yemeni, Egyptian, and Levantine. The dataset is split into training, validation, and testing sets to enhance its usage. The validation and testing sets contain 10 hours of audio segments each. Besides giving a detailed description of the dataset, wide range of speech recognition experiments using standard tools are also presented.
BibTeX
@inproceedings{icassp2024_sadasaudiaudioda,
title = {SADA: Saudi Audio Dataset for Arabic},
author = {Sadeen Alharbi and Areeb Alowisheq and Zoltán Tüske and Kareem Darwish and Abdullah Alrajeh and Abdulmajeed Alrowithi and Aljawharah Bin Tamran and Asma Ibrahim and Raghad Aloraini and Raneem Alnajim and Ranya Alkahtani and Renad Almuasaad and Sara Alrasheed and Shaykhah Alsubaie and Yaser Alonaizan},
booktitle = {ICASSP 2024},
year = {2024}
}