← Search

Reut Tsarfaty

27 accepted papers

2025

Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs

NAACL 2025findings

Despite advances in the multilingual capabilities of Large Language Models (LLMs) across diverse tasks, English remains the dominant language for LLM research and development. So, when working with a different language, this has led to the widespread practice of pre-translation, i.e., translating th…

2025

Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization

ACL 2025long

Automatic N-gram based metrics such as ROUGE are widely used for evaluating generative tasks such as summarization. While these metrics are considered indicative (even if imperfect), of human evaluation for English, their suitability for other languages remains unclear. To address this, in this pape…

2025

Superlatives in Context: Modeling the Implicit Semantics of Superlatives

NAACL 2025long

Superlatives are used to single out elements with a maximal/minimal property. Semantically, superlatives perform a set comparison: something (or some things) has the min/max property out of a set. As such, superlatives provide an ideal phenomenon for studying implicit phenomena and discourse restric…

2024

Breaking the Language Barrier: Can Direct Inference Outperform Pre-Translation in Multilingual LLM Applications?

NAACL 2024short

Large language models hold significant promise in multilingual applications. However, inherent biases stemming from predominantly English-centric pre-training have led to the widespread practice of pre-translation, i.e., translating non-English inputs to English before inference, leading to complexi…

Cited by 13SourcePDFScholar
2024

HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew

ACL 2024findings

While large language models (LLMs) excel in various natural language tasks in English, their performance in low-resource languages like Hebrew, especially for generative tasks such as abstractive summarization, remains unclear. The high morphological richness in Hebrew adds further challenges due to…

2024

Into the Unknown: Generating Geospatial Descriptions for New Environments

ACL 2024findings

Similar to vision-and-language navigation (VLN) tasks that focus on bridging the gap between vision and language for embodied navigation, the new Rendezvous (RVS) task requires reasoning over allocentric spatial relationships using non-sequential navigation instructions and maps. However, performanc…

Cited by 1SourcePDFScholar
2024

Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP

EMNLP 2024main

Improvements in language models’ capabilities have pushed their applications towards longer contexts, making long-context evaluation and development an active research area. However, many disparate use-cases are grouped together under the umbrella term of “long-context”, defined simply by the total…

Cited by 13SourcePDFScholar
2024

Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)

ACL 2024findings

Large Vision-Language Models (LVLMs) are an extension of Large Language Models (LLMs) that facilitate processing both image and text inputs, expanding AI capabilities. However, LVLMs struggle with object hallucinations due to their reliance on text cues and learned object co-occurrence biases. While…

Cited by 4SourcePDFScholar
2024

Multilingual Instruction Tuning With Just a Pinch of Multilinguality

ACL 2024findings

As instruction-tuned large language models (LLMs) gain global adoption, their ability to follow instructions in multiple languages becomes increasingly crucial. In this work, we investigate how multilinguality during instruction tuning of a multilingual LLM affects instruction-following across langu…

Cited by 31SourcePDFScholar
2024

Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance

ACL 2024findings

Despite it being the cornerstone of BPE, the most common tokenization algorithm, the importance of compression in the tokenization process is still unclear. In this paper, we argue for the theoretical importance of compression, that can be viewed as 0-gram language modeling where equal probability i…

Cited by 16SourcePDFScholar
2023

COHESENTIA: A Novel Benchmark of Incremental versus Holistic Assessment of Coherence in Generated Texts

EMNLP 2023long main

Coherence is a linguistic term that refers to the relations between small textual units (sentences, propositions), which make the text logically consistent and meaningful to the reader. With the advances of generative foundational models in NLP, there is a pressing need to automatically assess the h…

Cited by 0SourceScholar
2023

Covering Uncommon Ground: Gap-Focused Question Generation for Answer Assessment

ACL 2023short

Human communication often involves information gaps between the interlocutors. For example, in an educational dialogue a student often provides an answer that is incomplete, and there is a gap between this answer and the perfect one expected by the teacher. Successful dialogue then hinges on the tea…

Cited by 2SourcePDFScholar
2023

HeGeL: A Novel Dataset for Geo-Location from Hebrew Text

ACL 2023findings

The task of textual geolocation — retrieving the coordinates of a place based on a free-form language description — calls for not only grounding but also natural language understanding and geospatial reasoning. Even though there are quite a few datasets in English used for geolocation, they are curr…

2023

HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark

EMNLP 2023long findings

Current benchmarks for Hebrew Natural Language Processing (NLP) focus mainly on morpho-syntactic tasks, neglecting the semantic dimension of language understanding. To bridge this gap, we set out to deliver a Hebrew Machine Reading Comprehension (MRC) dataset, where MRC is to be realized as extract…

Cited by 0SourceScholar
2023

Is Probing All You Need? Indicator Tasks as an Alternative to Probing Embedding Spaces

EMNLP 2023long findings

The ability to identify and control different kinds of linguistic information encoded in vector representations of words has many use cases, especially for explainability and bias removal. This is usually done via a set of simple classification tasks, termed \textit{probes}, to evaluate the informat…

Cited by 0SourceScholar
2023

Multilingual Sequence-to-Sequence Models for Hebrew NLP

ACL 2023findings

Recent work attributes progress in NLP to large language models (LMs) with increased model size and large quantities of pretraining data. Despite this, current state-of-the-art LMs for Hebrew are both under-parameterized and under-trained compared to LMs in other languages. Additionally, previous wo…

Cited by 4SourcePDFScholar
2023

The Truth, The Whole Truth, and Nothing but the Truth: A New Benchmark Dataset for Hebrew Text Credibility Assessment

EMNLP 2023long findings

In the age of information overload, it is more important than ever to discern fact from fiction. From the internet to traditional media, we are constantly confronted with a deluge of information, much of which comes from politicians and other public figures who wield significant influence. In this p…

Cited by 0SourceScholar
2022

(Un)solving Morphological Inflection: Lemma Overlap Artificially Inflates Models’ Performance

ACL 2022short

In the domain of Morphology, Inflection is a fundamental and important task that gained a lot of traction in recent years, mostly via SIGMORPHON’s shared-tasks. With average accuracy above 0.9 over the scores of all languages, the task is considered mostly solved using relatively generic neural seq2…

2022

AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence Level

ACL 2022long

Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advances. While advances reported for English using PLMs are unprecedented, reported advances using PLMs for Hebrew are few and…

2022

Breakpoint Transformers for Modeling and Tracking Intermediate Beliefs

EMNLP 2022main

Can we teach models designed for language understanding tasks to track and improve their beliefs through intermediate points in text? Besides making their inner workings more transparent, this would also help make models more reliable and consistent. To this end, we propose a representation learning…

2022

Morphological Reinflection with Multiple Arguments: An Extended Annotation schema and a Georgian Case Study

ACL 2022short

In recent years, a flurry of morphological datasets had emerged, most notably UniMorph, aa multi-lingual repository of inflection tables. However, the flat structure of the current morphological annotation makes the treatment of some languages quirky, if not impossible, specifically in cases of poly…

Cited by 9SourcePDFScholar
2021

Asking It All: Generating Contextualized Questions for any Semantic Role

EMNLP 2021main

Asking questions about a situation is an inherent step towards understanding it. To this end, we introduce the task of role question generation, which, given a predicate mention and a passage, requires producing a set of questions asking about all possible semantic roles of the predicate. We develop…

2021

The Possible, the Plausible, and the Desirable: Event-Based Modality Detection for Language Processing

ACL 2021long

Modality is the linguistic ability to describe vents with added information such as how desirable, plausible, or feasible they are. Modality is important for many NLP downstream tasks such as the detection of hedging, uncertainty, speculation, and more. Previous studies that address modality detecti…