← Search

Ehud Reiter

11 accepted papers

2025

Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers’ Views and Experience of Reproducibility

EMNLP 2025

Over the past 10 years in NLP/ML, as in other fields of science, there has been growing interest in, and work on, reproducibility and methods for improving it. Identical experiments producing different results can be due to variation between samples of evaluation items or evaluators, but it can also

Cited by 0SourcePDFScholar
2025

SPHERE: An Evaluation Card for Human-AI Systems

ACL 2025finding

In the era of Large Language Models (LLMs), establishing effective evaluation methods and standards for diverse human-AI interaction systems is increasingly challenging. To encourage more transparent documentation and facilitate discussion on human-AI system evaluation design options, we present an…

2025

Scalability of Bayesian Network Structure Elicitation with Large Language Models: a Novel Methodology and Comparative Analysis

COLING 2025main

In this work, we propose a novel method for Bayesian Networks (BNs) structure elicitation that is based on the initialization of several LLMs with different experiences, independently querying them to create a structure of the BN, and further obtaining the final structure by majority voting. We comp…

Cited by 3SourcePDFScholar
2024

Ask the experts: sourcing a high-quality nutrition counseling dataset through Human-AI collaboration

EMNLP 2024finding

Large Language Models (LLMs) are being employed by end-users for various tasks, including sensitive ones such as health counseling, disregarding potential safety concerns. It is thus necessary to understand how adequately LLMs perform in such domains. We conduct a case study on ChatGPT in nutrition…

2024

Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo

NAACL 2024long

Neural Table-to-Text models tend to hallucinate, producing texts that contain factual errors. We investigate whether such errors in the output can be traced back to problems with the input. We manually annotated 1,837 texts generated by multiple models in the politics domain of the ToTTo dataset. We…

2023

Are Experts Needed? On Human Evaluation of Counselling Reflection Generation

ACL 2023long

Reflection is a crucial counselling skill where the therapist conveys to the client their interpretation of what the client said. Language models have recently been used to generate reflections automatically, but human evaluation is challenging, particularly due to the cost of hiring experts. Laypeo…

Cited by 14SourcePDFScholar
2023

Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP

ACL 2023findings

Human evaluation is widely regarded as the litmus test of quality in NLP. A basic requirementof all evaluations, but in particular where they are used for meta-evaluation, is that they should support the same conclusions if repeated. However, the reproducibility of human evaluations is virtually nev…

Cited by 24SourcePDFScholar
2022

Anno-MI: A Dataset of Expert-Annotated Counselling Dialogues

ICASSP 2022accepted

Research on natural language processing for counselling dialogue analysis has seen substantial development in recent years, but access to this area remains extremely limited due to the lack of publicly available expert-annotated therapy conversations. In this work, we introduce AnnoMI, the first pub…

Cited by 0SourceScholar
2022

Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation

EMNLP 2022industry

Evaluating automatically generated text is generally hard due to the inherently subjective nature of many aspects of the output quality. This difficulty is compounded in automatic consultation note generation by differing opinions between medical experts both about which patient statements should be…

Cited by 7SourcePDFScholar
2022

Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation

ACL 2022long

In recent years, machine learning models have rapidly become better at generating clinical consultation notes; yet, there is little work on how to properly evaluate the generated consultation notes to understand the impact they may have on both the clinician using them and the patient’s clinical saf…

2022

User-Driven Research of Medical Note Generation Software

NAACL 2022long

A growing body of work uses Natural Language Processing (NLP) methods to automatically generate medical notes from audio recordings of doctor-patient consultations. However, there are very few studies on how such systems could be used in clinical practice, how clinicians would adjust to using them,…

Cited by 22SourcePDFScholar