2025
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
EMNLP 2025
The increasing demand for domain-specific evaluation of large language models (LLMs) has led to the development of numerous benchmarks. These efforts often adhere to the principle of data scaling, relying on large corpora or extensive question-answer (QA) sets to ensure broad coverage. However, the