← Search

Alena Fenogenova

8 accepted papers

2025

2Columns1Row: A Russian Benchmark for Textual and Multimodal Table Understanding and Reasoning

EMNLP 2025

Table understanding is a crucial task in document processing and is commonly encountered in practical applications. We introduce 2Columns1Row, the first open-source benchmark for the table question answering task in Russian. This benchmark evaluates the ability of models to reason about the relation

Cited by 0SourcePDFScholar
2025

MMTEB: Massive Multilingual Text Embedding Benchmark

ICLR 2025poster

Text embeddings are typically evaluated on a narrow set of tasks, limited in terms of languages, domains, and task types. To circumvent this limitation and to provide a more comprehensive evaluation, we introduce the Massive Multilingual Text Embedding Benchmark (MMTEB) -- a large-scale community-dr…

2025

The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design

NAACL 2025long

Embedding models play a crucial role in Natural Language Processing (NLP) by creating text embeddings used in various tasks such as information retrieval and assessing semantic text similarity. This paper focuses on research related to embedding models in the Russian language. It introduces a new Ru…

2024

A Family of Pretrained Transformer Language Models for Russian

COLING 2024main

Transformer language models (LMs) are fundamental to NLP research methodologies and applications in various languages. However, developing such models specifically for the Russian language has received little attention. This paper introduces a collection of 13 Russian Transformer LMs, which spans en…

2024

MERA: A Comprehensive LLM Evaluation in Russian

ACL 2024long

Over the past few years, one of the most notable advancements in AI research has been in foundation models (FMs), headlined by the rise of language models (LMs). However, despite researchers’ attention and the rapid growth in LM application, the capabilities, limitations, and associated risks still…

2024

RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs

EMNLP 2024main

Minimal pairs are a well-established approach to evaluating the grammatical knowledge of language models. However, existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. This paper introduces the Russian Benchmark of…

2022

TAPE: Assessing Few-shot Russian Language Understanding

EMNLP 2022finding

Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes. However, this fast-growing area lacks standardized evaluation suites for non-English languages, hindering progress outside the Anglo-centric paradigm. To address this line of researc…

2020

Read and Reason with MuSeRC and RuCoS: Datasets for Machine Reading Comprehension for Russian

COLING 2020main

The paper introduces two Russian machine reading comprehension (MRC) datasets, called MuSeRC and RuCoS, which require reasoning over multiple sentences and commonsense knowledge to infer the answer. The former follows the design of MultiRC, while the latter is a counterpart of the ReCoRD dataset. Th…