← Search

Rustem Yeshpanov

4 accepted papers

2024

KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis

COLING 2024main

This study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs, with a total duration of 74.85 hours, featuring 34.23 hours delivered by a female narrator and 40.62 hours by two male n…

2024

KazParC: Kazakh Parallel Corpus for Machine Translation

COLING 2024main

We introduce KazParC, a parallel corpus designed for machine translation across Kazakh, English, Russian, and Turkish. The first and largest publicly available corpus of its kind, KazParC contains a collection of 371,902 parallel sentences covering different domains and developed with the assistance…

2024

KazQAD: Kazakh Open-Domain Question Answering Dataset

COLING 2024main

We introduce KazQAD—a Kazakh open-domain question answering (ODQA) dataset—that can be used in both reading comprehension and full ODQA settings, as well as for information retrieval experiments. KazQAD contains just under 6,000 unique questions with extracted short answers and nearly 12,000 passage…

2024

KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and Attitudes

COLING 2024main

This paper presents KazSAnDRA, a dataset developed for Kazakh sentiment analysis that is the first and largest publicly available dataset of its kind. KazSAnDRA comprises an extensive collection of 180,064 reviews obtained from various sources and includes numerical ratings ranging from 1 to 5, prov…