← Search

Matúš Pikuliak

4 accepted papers

2024

Disinformation Capabilities of Large Language Models

ACL 2024long

Automated disinformation generation is often listed as one of the risks of large language models (LLMs). The theoretical ability to flood the information space with disinformation content might have dramatic consequences for democratic societies around the world. This paper presents a comprehensive…

2024

Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling

EMNLP 2024finding

We present GEST – a new manually created dataset designed to measure gender-stereotypical reasoning in language models and machine translation systems. GEST contains samples for 16 gender stereotypes about men and women (e.g., Women are beautiful, Men are leaders) that are compatible with the Englis…

2023

MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark

EMNLP 2023long main

There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also reflected in the available benchmarks which lack authentic texts in languages ot…

Cited by 0SourcecodeScholar
2022

SlovakBERT: Slovak Masked Language Model

EMNLP 2022finding

We introduce a new Slovak masked language model called SlovakBERT. This is to our best knowledge the first paper discussing Slovak transformers-based language models. We evaluate our model on several NLP tasks and achieve state-of-the-art results. This evaluation is likewise the first attempt to est…