← Search

Ekaterina Artemova

12 accepted papers

2025

Beemo: Benchmark of Expert-edited Machine-generated Outputs

NAACL 2025long

The rapid proliferation of large language models (LLMs) has increased the volume of machine-generated texts (MGTs) and blurred text authorship in various domains. However, most existing MGT benchmarks include single-author texts (human-written and machine-generated). This conventional design fails t…

2024

LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection

EMNLP 2024system demonstrations

The ease of access to large language models (LLMs) has enabled a widespread of machine-generated texts, and now it is often hard to tell whether a piece of text was human-written or machine-generated. This raises concerns about potential misuse, particularly within educational and academic domains.…

2024

RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs

EMNLP 2024main

Minimal pairs are a well-established approach to evaluating the grammatical knowledge of language models. However, existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. This paper introduces the Russian Benchmark of…

2024

RuBia: A Russian Language Bias Detection Dataset

COLING 2024main

Warning: this work contains upsetting or disturbing content. Large language models (LLMs) tend to learn the social and cultural biases present in the raw pre-training data. To test if an LLM’s behavior is fair, functional datasets are employed, and due to their purpose, these datasets are highly lan…

2024

Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data

COLING 2024main

Named Entity Recognition (NER) is a fundamental task to extract key information from texts, but annotated resources are scarce for dialects. This paper introduces the first dialectal NER dataset for German, BarNER, with 161K tokens annotated on Bavarian Wikipedia articles (bar-wiki) and tweets (bar-…

2023

Boosting Zero-shot Cross-lingual Retrieval by Training on Artificially Code-Switched Data

ACL 2023findings

Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in diff…

2022

Acceptability Judgements via Examining the Topology of Attention Maps

EMNLP 2022finding

The role of the attention mechanism in encoding linguistic knowledge has received special interest in NLP. However, the ability of the attention heads to judge the grammatical acceptability of a sentence has been underexplored. This paper approaches the paradigm of acceptability judgments with topol…

2022

RuCoLA: Russian Corpus of Linguistic Acceptability

EMNLP 2022main

Linguistic acceptability (LA) attracts the attention of the research community due to its many uses, such as testing the grammatical knowledge of language models and filtering implausible texts with acceptability classifiers.However, the application scope of LA in languages other than English is lim…

2022

TAPE: Assessing Few-shot Russian Language Understanding

EMNLP 2022finding

Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes. However, this fast-growing area lacks standardized evaluation suites for non-English languages, hindering progress outside the Anglo-centric paradigm. To address this line of researc…

2021

Artificial Text Detection via Examining the Topology of Attention Maps

EMNLP 2021main

The impressive capabilities of recent generative models to create texts that are challenging to distinguish from the human-written ones can be misused for generating fake news, product reviews, and even abusive content. Despite the prominent performance of existing methods for artificial text detect…

2021

Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection

AAAI 2021technical

Real-life applications, heavily relying on machine learning, such as dialog systems, demand for out-of-domain detection methods. Intent classification models should be equipped with a mechanism to distinguish seen intents from unseen ones so that the dialog agent is capable of rejecting the latter a…

2020

SumTitles: a Summarization Dataset with Low Extractiveness

COLING 2020main

The existing dialogue summarization corpora are significantly extractive. We introduce a methodology for dataset extractiveness evaluation and present a new low-extractive corpus of movie dialogues for abstractive text summarization along with baseline evaluation. The corpus contains 153k dialogues…