2025
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
ACL 2025long
Large Language Models (LLMs) drive scientific question-answering on modern search engines, yet their evaluation robustness remains underexplored. We introduce YESciEval, an open-source framework that combines fine-grained rubric-based assessment with reinforcement learning to mitigate optimism bias…