← Search

Marzena Karpinska

14 accepted papers

2025

An Interdisciplinary Approach to Human-Centered Machine Translation

EMNLP 2025

Machine Translation (MT) tools are widely used today, often in contexts where professional translators are not present. Despite progress in MT technology, a gap persists between system development and real-world usage, particularly for non-expert users who may struggle to assess translation reliabil

Cited by 0SourcePDFScholar
2025

Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code

COLING 2025industry

Pretrained language models are integral part of AI applications, but their high computational cost for training limits accessibility. Initiatives such as Bloom and StarCoder aim to democratize access to pretrained models for collaborative community development. Despite these efforts, such models enc…

Cited by 2SourcePDFScholar
2025

CaLMQA: Exploring culturally specific long-form question answering across 23 languages

ACL 2025long

Despite rising global usage of large language models (LLMs), their ability to generate *long-form* answers to *culturally specific* questions remains unexplored in many languages. To fill this gap, we perform the first study of textual multilingual long-form QA by creating CaLMQA, a dataset of **51.…

2025

Does quantization affect models’ performance on long-context tasks?

EMNLP 2025

Large language models (LLMs) now support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency. Quantization can mitigate these costs, but may degrade performance. In this work, we present the first systematic evaluation of quantized LL

2025

OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature

EMNLP 2025

Large language models (LLMs) are known to memorize and recall English text from their pretraining data. However, the extent to which this ability generalizes to non-English languages or transfers across languages remains unclear. This paper investigates multilingual and cross-lingual memorization in

2025

People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text

ACL 2025long

In this paper, we study how well humans can detect text generated by commercial LLMs (GPT-4o, Claude, o1). We hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions. Our experiments s…

2024

NarrativeTime: Dense Temporal Annotation on a Timeline

COLING 2024main

For the past decade, temporal annotation has been sparse: only a small portion of event pairs in a text was annotated. We present NarrativeTime, the first timeline-based annotation framework that achieves full coverage of all possible TLINKs. To compare with the previous SOTA in dense temporal annot…

2024

One Thousand and One Pairs: A “novel” challenge for long-context language models

EMNLP 2024main

Synthetic long-context LLM benchmarks (e.g., “needle-in-the-haystack”) test only surface-level retrieval capabilities; but how well can long-context LLMs retrieve, synthesize, and reason over information across book-length inputs? We address this question by creating NoCha, a dataset of 1,001 minima…

2023

Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense

NeurIPS 2023poster

The rise in malicious usage of large language models, such as fake content creation and academic plagiarism, has motivated the development of approaches that identify AI-generated text, including those based on watermarking or outlier detection. However, the robustness of these detection algorithms…

2023

Program Chairs’ Report on Peer Review at ACL 2023

ACL 2023long

We present a summary of the efforts to improve conference peer review that were implemented at ACL’23. This includes work with the goal of improving review quality, clearer workflow and decision support for the area chairs, as well as our efforts to improve paper-reviewer matching for various kinds…

2022

DEMETR: Diagnosing Evaluation Metrics for Translation

EMNLP 2022main

While machine translation evaluation metrics based on string overlap (e.g., BLEU) have their limitations, their computations are transparent: the BLEU score assigned to a particular candidate translation can be traced back to the presence or absence of certain words. The operations of newer learned…

2022

Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature

EMNLP 2022main

Literary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators relative to the many untranslated works published around the world. Machine translation (MT) holds potential to complement the work of human translators by improving bo…

2022

Revisiting Statistical Laws of Semantic Shift in Romance Cognates

COLING 2022main

This article revisits statistical relationships across Romance cognates between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy. Cognates are words that are derived from a common etymon, in this case, a Latin ancestor. Despite their shared etymology, some co…

Cited by 4SourcePDFScholar
2021

The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation

EMNLP 2021main

Recent text generation research has increasingly focused on open-ended domains such as story and poetry generation. Because models built for such tasks are difficult to evaluate automatically, most researchers in the space justify their modeling choices by collecting crowdsourced human judgments of…

Cited by 136SourcePDFScholar