← Search

Jeffrey Cheng

2 accepted papers

2025

CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?

EMNLP 2025

A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically generate plausible (if generic) reviews, ensuring that these reviews are sound and grounded in the papers’ claims remains chal

2025

Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering

ACL 2025short

Scaling the test-time compute of large language models has demonstrated impressive performance on reasoning benchmarks. However, existing evaluations of test-time scaling make the strong assumption that a reasoning system should always give an answer to any question provided. This overlooks concerns…