← Search

Roy Bar-Haim

6 accepted papers

2026

CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

AAAI 2026technical

The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specif

Cited by 0SourcePDFScholar
2025

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation

EMNLP 2025

We introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges. Evaluating debate speeches requires a deep understanding of the speech at multiple levels, including argument strength and relevance, the coherence and organization of the speech, the appropriateness

2025

JuStRank: Benchmarking LLM Judges for System Ranking

ACL 2025long

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the use of LLM-based judges a compelling solution for this challenge. Crucially, this…

Cited by 0SourcePDFScholar
2023

From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization

ACL 2023long

Key Point Analysis (KPA) has been recently proposed for deriving fine-grained insights from collections of textual comments. KPA extracts the main points in the data as a list of concise sentences or phrases, termed Key Points, and quantifies their prevalence. While key points are more expressive th…

2021

Every Bite Is an Experience: Key Point Analysis of Business Reviews

ACL 2021long

Previous work on review summarization focused on measuring the sentiment toward the main aspects of the reviewed product or business, or on creating a textual summary. These approaches provide only a partial view of the data: aspect-based sentiment summaries lack sufficient explanation or justificat…

Cited by 25SourcePDFScholar
2021

Project Debater APIs: Decomposing the AI Grand Challenge

EMNLP 2021system demonstrations

Project Debater was revealed in 2019 as the first AI system that can debate human experts on complex topics. Engaging in a live debate requires a diverse set of skills, and Project Debater has been developed accordingly as a collection of components, each designed to perform a specific subtask. Proj…