← Search

Sergey Pletenev

6 accepted papers

2025

Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home

ACL 2025long

Retrieval Augmented Generation (RAG) improves correctness of Question Answering (QA) and addresses hallucinations in Large Language Models (LLMs), yet greatly increase computational costs. Besides, RAG is not always needed as may introduce irrelevant information. Recent adaptive retrieval methods in…

2025

How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?

NAACL 2025findings

The performance of Large Language Models (LLMs) on many tasks is greatly limited by the knowledge learned during pre-training and stored in the model’s parameters. Low-rank adaptation (LoRA) is a popular and efficient training technique for updating or domain-specific adaptation of LLMs. In this stu…

2025

LLM-Independent Adaptive RAG: Let the Question Speak for Itself

EMNLP 2025

Large Language Models (LLMs) are prone to hallucinations, and Retrieval-Augmented Generation (RAG) helps mitigate this, but at a high computational cost while risking misinformation. Adaptive retrieval aims to retrieve only when necessary, but existing approaches rely on LLM-based uncertainty estima

2025

SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators

NAACL 2025long

Existing approaches to multilingual text detoxification are hampered by the scarcity of parallel multilingual datasets. In this work, we introduce a pipeline for the generation of multilingual parallel detoxification data. We also introduce SynthDetoxM, a manually collected and synthetically generat…

2025

Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

EMNLP 2025

Large Language Models (LLMs) often hallucinate in question answering (QA) tasks. A key yet underexplored factor contributing to this is the temporality of questions – whether they are evergreen (answers remain stable over time) or mutable (answers change). In this work, we introduce EverGreenQA, the

2024

LLMs to Replace Crowdsourcing For Parallel Data Creation? The Case of Text Detoxification

EMNLP 2024finding

The lack of high-quality training data remains a significant challenge in NLP. Manual annotation methods, such as crowdsourcing, are costly, require intricate task design skills, and, if used incorrectly, may result in poor data quality. From the other hand, LLMs have demonstrated proficiency in man…