← Search

Ekaterina Kochmar

13 accepted papers

2025

A Fully Automated Pipeline for Conversational Discourse Annotation: Tree Scheme Generation and Labeling with Large Language Models

ACL 2025finding

Recent advances in Large Language Models (LLMs) have shown promise in automating discourse annotation for conversations. While manually designing tree annotation schemes significantly improves annotation quality for humans and models, their creation remains time-consuming and requires expert knowled…

2025

KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan

ACL 2025long

Despite having a population of twenty million, Kazakhstan’s culture and language remain underrepresented in the field of natural language processing. Although large language models (LLMs) continue to advance worldwide, progress in Kazakh language has been limited, as seen in the scarcity of dedicate…

Cited by 0SourcePDFScholar
2025

LLMs cannot spot math errors, even when allowed to peek into the solution

EMNLP 2025

Large language models (LLMs) demonstrate remarkable performance on math word problems, yet they have been shown to struggle with meta-reasoning tasks such as identifying errors in student solutions. In this work, we investigate the challenge of locating the first error step in stepwise solutions usi

Cited by 0SourcePDFScholar
2025

SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models

ACL 2025finding

Large language models (LLMs) have been widely adopted due to their remarkable performance across various applications, driving the accelerated development of a large number of diverse models. However, these individual LLMs show limitations in generalization and performance on complex tasks due to in…

2025

Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors

NAACL 2025long

In this paper, we investigate whether current state-of-the-art large language models (LLMs) are effective as AI tutors and whether they demonstrate pedagogical abilities necessary for good AI tutoring in educational dialogues. Previous efforts towards evaluation have beenlimited to subjective protoc…

2025

UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment

EMNLP 2025

We introduce UniversalCEFR, a large-scale multilingual multidimensional dataset of texts annotated according to the CEFR (Common European Framework of Reference) scale in 13 languages. To enable open research in both automated readability and language proficiency assessment, UniversalCEFR comprises

2025

What Makes Cryptic Crosswords Challenging for LLMs?

COLING 2025main

Cryptic crosswords are puzzles that rely on general knowledge and the solver’s ability to manipulate language on different levels, dealing with various types of wordplay. Previous research suggests that solving such puzzles is challenging even for modern NLP models, including Large Language Models (…

2024

How Teachers Can Use Large Language Models and Bloom’s Taxonomy to Create Educational Quizzes

AAAI 2024technical

Question generation (QG) is a natural language processing task with an abundance of potential benefits and use cases in the educational domain. In order for this potential to be realized, QG systems must be designed and validated with pedagogical needs in mind. However, little research has assessed…

Cited by 16SourcePDFScholar
2023

Automatic Readability Assessment for Closely Related Languages

ACL 2023findings

In recent years, the main focus of research on automatic readability assessment (ARA) has shifted towards using expensive deep learning-based methods with the primary goal of increasing models’ accuracy. This, however, is rarely applicable for low-resource languages where traditional handcrafted fea…

2023

BasahaCorpus: An Expanded Linguistic Resource for Readability Assessment in Central Philippine Languages

EMNLP 2023short main

Current research on automatic readability assessment (ARA) has focused on improving the performance of models in high-resource languages such as English. In this work, we introduce and release BasahaCorpus as part of an initiative aimed at expanding available corpora and baseline models for readabil…

Cited by 0SourcecodeScholar
2021

Word Complexity is in the Eye of the Beholder

NAACL 2021long

Lexical complexity is a highly subjective notion, yet this factor is often neglected in lexical simplification and readability systems which use a ”one-size-fits-all” approach. In this paper, we investigate which aspects contribute to the notion of lexical complexity in various groups of readers, fo…

Cited by 19SourcePDFScholar