← Search

Jakub Simko

9 accepted papers

2025

A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages

EMNLP 2025

Large Language Models (LLMs) are increasingly used to generate synthetic textual data for training smaller specialized models. However, a comparison of various generation strategies for low-resource language settings is lacking. While various prompting strategies have been proposed—such as demonstra

Cited by 0SourcePDFScholar
2025

LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?

NAACL 2025long

The generative large language models (LLMs) are increasingly being used for data augmentation tasks, where text samples are LLM-paraphrased and then used for classifier fine-tuning. Previous studies have compared LLM-based augmentations with established augmentation techniques, but the results are c…

2025

Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation

EMNLP 2025

The generative large language models (LLMs) are increasingly used for data augmentation tasks, where text samples are paraphrased (or generated anew) and then used for downstream model fine-tuning. This is useful, especially for low-resource settings. For better augmentations, LLMs are prompted with

Cited by 0SourcePDFScholar
2024

Authorship Obfuscation in Multilingual Machine-Generated Text Detection

EMNLP 2024finding

High-quality text generation capability of latest Large Language Models (LLMs) causes concerns about their misuse (e.g., in massive generation/spread of disinformation). Machine-generated text (MGT) detection is important to cope with such threats. However, it is susceptible to authorship obfuscatio…

2024

Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation

ACL 2024long

The latest generative large language models (LLMs) have found their application in data augmentation tasks, where small numbers of text samples are LLM-paraphrased and then used to fine-tune downstream models. However, more research is needed to assess how different prompts, seed data selection stra…

2024

Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation

EMNLP 2024finding

While fine-tuning of pre-trained language models generally helps to overcome the lack of labelled training samples, it also displays model performance instability. This instability mainly originates from randomness in initialisation or data shuffling. To address this, researchers either modify the t…

2023

ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness

EMNLP 2023long main

The emergence of generative large language models (LLMs) raises the question: what will be its impact on crowdsourcing? Traditionally, crowdsourcing has been used for acquiring solutions to a wide variety of human-intelligence tasks, including ones involving text generation, modification or evaluati…

Cited by 68SourceScholar
2023

MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark

EMNLP 2023long main

There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also reflected in the available benchmarks which lack authentic texts in languages ot…

Cited by 0SourcecodeScholar
2022

Black-box Audit of YouTube's Video Recommendation: Investigation of Misinformation Filter Bubble Dynamics (Extended Abstract)

IJCAI 2022poster

In this paper, we describe a black-box sockpuppeting audit which we carried out to investigate the creation and bursting dynamics of misinformation filter bubbles on YouTube. Pre-programmed agents acting as YouTube users stimulated YouTube's recommender systems: they first watched a series of misinf…

Cited by 0SourcePDFScholar