COLING 2024main2 citations

SlovakSum: A Large Scale Slovak Summarization Dataset

Viktoria Ondrejova, Marek Suppa

Abstract

The ability to automatically summarize news articles has become increasingly important due to the vast amount of information available online. Together with the rise of chatbots , Natural Language Processing (NLP) has recently experienced a tremendous amount of development. Despite these advancements, the majority of research is focused on established well-resourced languages, such as English. To contribute to development of the low resource Slovak language, we introduce SlovakSum, a Slovak news summarization dataset consisting of over 200 thousand news articles with titles and short abstracts obtained from multiple Slovak newspapers. The abstractive approach, including MBART and mT5 models, was used to evaluate various baselines. The code for the reproduction of our dataset and experiments can be found at https://github.com/NaiveNeuron/slovaksum

BibTeX
@inproceedings{ondrejova-suppa-2024-slovaksum,
    title = "{S}lovak{S}um: A Large Scale {S}lovak Summarization Dataset",
    author = "Ondrejova, Viktoria  and
      Suppa, Marek",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.1298/",
    pages = "14916--14922"
}
SlovakSum: A Large Scale Slovak Summarization Dataset · COLING 2024