← Search

Lucia Urcelay Ganzabal

1 accepted papers

2025

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

NAACL 2025short

Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. Close-ended measurements evaluate the factuality of responses but lack expressiveness. Open-ended capture the model’s capacity to produce discourse re…

Cited by 1SourcePDFScholar