← Search

Nitay Calderon

13 accepted papers

2026

Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality

ICML 2026poster

Standard factuality evaluations of LLMs treat all errors alike, obscuring whether failures arise from missing knowledge (empty shelves) or from limited access to encoded facts (lost keys). We propose a behavioral framework that profiles factual knowledge at the level of facts rather than questions, …

Cited by 0SourceScholar
2025

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance

EMNLP 2025

NLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field. Traditionally, expert annotations ensure high-quality labels; however, the cost of expert annotation does not scale well with the growing demand for larger datasets required by mo

Cited by 0SourcePDFScholar
2025

Dementia Through Different Eyes: Explainable Modeling of Human and LLM Perceptions for Early Awareness

EMNLP 2025

Cognitive decline often surfaces in language years before diagnosis. It is frequently non-experts, such as those closest to the patient, who first sense a change and raise concern. As LLMs become integrated into daily communication and used over prolonged periods, it may even be an LLM that notices

Cited by 0SourcePDFScholar
2025

NL-Eye: Abductive NLI For Images

ICLR 2025poster

Will a Visual Language Model (VLM)-based bot warn us about slipping if it detects a wet floor? Recent VLMs have demonstrated impressive capabilities, yet their ability to infer outcomes and causes remains underexplored. To address this, we introduce NL-Eye, a benchmark designed to assess VLMs' visua…

Cited by 0SourcePDFScholar
2025

On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs

NAACL 2025long

Recent advancements in NLP systems, particularly with the introduction of LLMs, have led to widespread adoption of these systems by a broad spectrum of users across various domains, impacting decision-making, the job market, society, and scientific research. This surge in usage has led to an explosi…

2025

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

ACL 2025long

The “LLM-as-an-annotator” and “LLM-as-a-judge” paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans. LLM annotations are widely used, not only in NLP research but also in fields like medicine, psychology, and social science. D…

2024

Faithful Explanations of Black-box NLP Models Using LLM-generated Counterfactuals

ICLR 2024poster

Causal explanations of the predictions of NLP systems are essential to ensure safety and establish trust. Yet, existing methods often fall short of explaining model predictions effectively or efficiently and are often model-specific. In this paper, we address model-agnostic explanations, proposing t…

Cited by 41SourcePDFScholar
2024

Measuring the Robustness of NLP Models to Domain Shifts

EMNLP 2024finding

Existing research on Domain Robustness (DR) suffers from disparate setups, limited task variety, and scarce research on recent capabilities such as in-context learning. Furthermore, the common practice of measuring DR might not be fully accurate. Current research focuses on challenge sets and relies…

2024

The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth

NAACL 2024long

Queer youth face increased mental health risks, such as depression, anxiety, and suicidal ideation. Hindered by negative stigma, they often avoid seeking help and rely on online resources, which may provide incompatible information. Although access to a supportive environment and reliable informatio…

2023

A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target Training

ACL 2023long

Modern Natural Language Generation (NLG) models come with massive computational and storage requirements. In this work, we study the potential of compressing them, which is crucial for real-world applications serving millions of users. We focus on Knowledge Distillation (KD) techniques, in which a s…

2022

A Functional Information Perspective on Model Interpretation

ICML 2022spotlight

Contemporary predictive models are hard to interpret as their deep nets exploit numerous complex relations between input elements. This work suggests a theoretical framework for model interpretability by measuring the contribution of relevant features to the functional entropy of the network with re…

2022

DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation

ACL 2022long

Natural language processing (NLP) algorithms have become very successful, but they still struggle when applied to out-of-distribution examples. In this paper we propose a controllable generation approach in order to deal with this domain adaptation (DA) challenge. Given an input text example, our Do…