← Search

Xiechi Zhang

4 accepted papers

2025

ACE-M3: Automatic Capability Evaluator for Multimodal Medical Models

COLING 2025main

As multimodal large language models (MLLMs) gain prominence in the medical field, the need for precise evaluation methods to assess their effectiveness has become critical. While benchmarks provide a reliable means to evaluate the capabilities of MLLMs, traditional metrics like ROUGE and BLEU employ…

Cited by 0SourcePDFScholar
2025

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation

ACL 2025long

With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, which rely on token overlaps to measure quality, significantly overlook the import…

Cited by 0SourcePDFScholar
2025

Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation

AAAI 2025technical

In the rapidly evolving landscape of large language models (LLMs) for medical applications, ensuring the reliability and accuracy of these models in clinical settings is paramount. Existing benchmarks often focus on fixed-format tasks like multiple-choice QA, which fail to capture the complexity of…

2023

A Disentangled-Attention Based Framework with Persona-Aware Prompt Learning for Dialogue Generation

AAAI 2023technical

Endowing dialogue agents with personas is the key to delivering more human-like conversations. However, existing persona-grounded dialogue systems still lack informative details of human conversations and tend to reply with inconsistent and generic responses. One of the main underlying causes is tha…

Cited by 5SourcePDFScholar