← Search

Peter Rupnik

2 accepted papers

2024

Do Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages

COLING 2024main

Large, curated, web-crawled corpora play a vital role in training language models (LMs). They form the lion’s share of the training data in virtually all recent LMs, such as the well-known GPT, LLaMA and XLM-RoBERTa models. However, despite this importance, relatively little attention has been given…

2024

The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings

COLING 2024main

The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which are used in a series of experiments focused on training a robust sentiment identifier for parliamentary proceedings. The paper additionally introduces the first domain-specific multilingual…

Cited by 8SourcePDFScholar