2025
Can Large Language Models Accurately Generate Answer Keys for Health-related Questions?
ACL 2025short
The evaluation of text generated by LLMs remains a challenge for question answering, retrieval augmented generation (RAG), summarization, and many other natural language processing tasks. Evaluating the factuality of LLM generated responses is particularly important in medical question answering, wh…