← Search

Branislav Pecher

7 accepted papers

2025

Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance

EMNLP 2025

When solving NLP tasks with limited labelled data, researchers typically either use a general large language model without further update, or use a small number of labelled samples to tune a specialised smaller model. In this work, we answer an important question – how many labelled samples are requ

Cited by 0SourcePDFScholar
2025

Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation

EMNLP 2025

The generative large language models (LLMs) are increasingly used for data augmentation tasks, where text samples are paraphrased (or generated anew) and then used for downstream model fine-tuning. This is useful, especially for low-resource settings. For better augmentations, LLMs are prompted with

Cited by 0SourcePDFScholar
2024

Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation

ACL 2024long

The latest generative large language models (LLMs) have found their application in data augmentation tasks, where small numbers of text samples are LLM-paraphrased and then used to fine-tune downstream models. However, more research is needed to assess how different prompts, seed data selection stra…

2024

Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation

EMNLP 2024finding

While fine-tuning of pre-trained language models generally helps to overcome the lack of labelled training samples, it also displays model performance instability. This instability mainly originates from randomness in initialisation or data shuffling. To address this, researchers either modify the t…

2024

On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices

EMNLP 2024main

While learning with limited labelled data can effectively deal with a lack of labels, it is also sensitive to the effects of uncontrolled randomness introduced by so-called randomness factors (i.e., non-deterministic decisions such as choice or order of samples). We propose and formalise a method to…

2022

Black-box Audit of YouTube's Video Recommendation: Investigation of Misinformation Filter Bubble Dynamics (Extended Abstract)

IJCAI 2022poster

In this paper, we describe a black-box sockpuppeting audit which we carried out to investigate the creation and bursting dynamics of misinformation filter bubbles on YouTube. Pre-programmed agents acting as YouTube users stimulated YouTube's recommender systems: they first watched a series of misinf…

Cited by 0SourcePDFScholar
2022

Transferability and Stability of Learning with Limited Labelled Data in Multilingual Text Document Classification

IJCAI 2022poster

We focus on learning with limited labelled data (especially meta-learning) in conjunction with so-far under-researched multilingual textual document classification. The core principle in such learning is to achieve transferability of learned knowledge to new datasets and tasks. Currently, factors in…

Cited by 0SourcePDFScholar