← Search

Gili Lior

2 accepted papers

2025

ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments

EMNLP 2025

LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments evaluation over the space of meaning-preserving prompt pertu

2024

Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction

ACL 2024findings

Document collections of various domains, e.g., legal, medical, or financial, often share some underlying collection-wide structure, which captures information that can aid both human users and structure-aware models.We propose to identify the typical structure of document within a collection, which…