2025
ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation
ACL 2025finding
Evaluating personalized text generated by large language models (LLMs) is challenging, as only the LLM user, i.e. prompt author, can reliably assess the output, but re-engaging the same individuals across studies is infeasible. This paper addresses the challenge of evaluating personalized text gener…