← Search

Jan Cegin

6 accepted papers

2025

A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages

EMNLP 2025

Large Language Models (LLMs) are increasingly used to generate synthetic textual data for training smaller specialized models. However, a comparison of various generation strategies for low-resource language settings is lacking. While various prompting strategies have been proposed—such as demonstra

Cited by 0SourcePDFScholar
2025

LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?

NAACL 2025long

The generative large language models (LLMs) are increasingly being used for data augmentation tasks, where text samples are LLM-paraphrased and then used for classifier fine-tuning. Previous studies have compared LLM-based augmentations with established augmentation techniques, but the results are c…

2025

Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation

EMNLP 2025

The generative large language models (LLMs) are increasingly used for data augmentation tasks, where text samples are paraphrased (or generated anew) and then used for downstream model fine-tuning. This is useful, especially for low-resource settings. For better augmentations, LLMs are prompted with

Cited by 0SourcePDFScholar
2024

Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation

ACL 2024long

The latest generative large language models (LLMs) have found their application in data augmentation tasks, where small numbers of text samples are LLM-paraphrased and then used to fine-tune downstream models. However, more research is needed to assess how different prompts, seed data selection stra…

2024

Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation

EMNLP 2024finding

While fine-tuning of pre-trained language models generally helps to overcome the lack of labelled training samples, it also displays model performance instability. This instability mainly originates from randomness in initialisation or data shuffling. To address this, researchers either modify the t…

2023

ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness

EMNLP 2023long main

The emergence of generative large language models (LLMs) raises the question: what will be its impact on crowdsourcing? Traditionally, crowdsourcing has been used for acquiring solutions to a wide variety of human-intelligence tasks, including ones involving text generation, modification or evaluati…

Cited by 68SourceScholar