← Search

Laida Kushnareva

5 accepted papers

2025

Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders

ACL 2025finding

Artificial Text Detection (ATD) is becoming increasingly important with the rise of advanced Large Language Models (LLMs). Despite numerous efforts, no single algorithm performs consistently well across different types of unseen text or guarantees effective generalization to new LLMs. Interpretabili…

2025

Quantifying Logical Consistency in Transformers via Query-Key Alignment

EMNLP 2025

Large language models (LLMs) excel at many NLP tasks, yet their multi-step logical reasoning remains unreliable. Existing solutions such as Chain-of-Thought prompting generate intermediate steps but provide no internal check of their logical coherence. In this paper, we use the “QK-score”, a lightwe

Cited by 0SourcePDFScholar
2024

Robust AI-Generated Text Detection by Restricted Embeddings

EMNLP 2024finding

Growing amount and quality of AI-generated texts makes detecting such content more difficult. In most real-world scenarios, the domain (style and topic) of generated data and the generator model are not known in advance. In this work, we focus on the robustness of classifier-based detectors of AI-ge…

2022

Acceptability Judgements via Examining the Topology of Attention Maps

EMNLP 2022finding

The role of the attention mechanism in encoding linguistic knowledge has received special interest in NLP. However, the ability of the attention heads to judge the grammatical acceptability of a sentence has been underexplored. This paper approaches the paradigm of acceptability judgments with topol…

2021

Artificial Text Detection via Examining the Topology of Attention Maps

EMNLP 2021main

The impressive capabilities of recent generative models to create texts that are challenging to distinguish from the human-written ones can be misused for generating fake news, product reviews, and even abusive content. Despite the prominent performance of existing methods for artificial text detect…