← Search

Lilach Eden

4 accepted papers

2026

CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

AAAI 2026technical

The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specif

Cited by 0SourcePDFScholar
2025

JuStRank: Benchmarking LLM Judges for System Ranking

ACL 2025long

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the use of LLM-based judges a compelling solution for this challenge. Crucially, this…

Cited by 0SourcePDFScholar
2023

From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization

ACL 2023long

Key Point Analysis (KPA) has been recently proposed for deriving fine-grained insights from collections of textual comments. KPA extracts the main points in the data as a list of concise sentences or phrases, termed Key Points, and quantifies their prevalence. While key points are more expressive th…

2021

Every Bite Is an Experience: Key Point Analysis of Business Reviews

ACL 2021long

Previous work on review summarization focused on measuring the sentiment toward the main aspects of the reviewed product or business, or on creating a textual summary. These approaches provide only a partial view of the data: aspect-based sentiment summaries lack sufficient explanation or justificat…

Cited by 25SourcePDFScholar