2025
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
ACL 2025long
With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, which rely on token overlaps to measure quality, significantly overlook the import…