← Search

Chacha Chen

6 accepted papers

2026

Collaborative Disagreement Resolution for Scalable Oversight

ICML 2026poster

*Debate*, where AI agents argue opposing positions, has emerged as a key approach to scalable oversight. However, debate faces a fundamental tension: models are incentivized to be persuasive to the judge, which may not always align with epistemic honesty. In this work, we propose an alternative para…

Cited by 0SourceScholar
2025

CLEAR: A Clinically Grounded Tabular Framework for Radiology Report Evaluation

EMNLP 2025

Existing metrics often lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports, resulting in suboptimal evaluation. We introduce a **Cl**inically grounded tabular framework with **E**xpert-curated labels and **A**ttribute

2025

GPT-4V Cannot Generate Radiology Reports Yet

NAACL 2025findings

GPT-4’s purported strong multimodal abilities raise interests in using it to automate radiology report writing, but there lacks thorough evaluations. In this work, we perform a systematic evaluation of GPT-4 (4o and vision-preview) in generating radiology reports across three chest X-ray report benc…

Cited by 3SourcePDFScholar
2023

Learning Human-Compatible Representations for Case-Based Decision Support

ICLR 2023poster

Algorithmic case-based decision support provides examples to help human make sense of predicted labels and aid human in decision-making tasks. Despite the promising performance of supervised learning, representations learned by supervised models may not align well with human intuitions: what models…

2022

Learning to Rank Visual Stories From Human Ranking Data

ACL 2022long

Visual storytelling (VIST) is a typical vision and language task that has seen extensive development in the natural language generation research domain. However, it remains unclear whether conventional automatic evaluation metrics for text generation are applicable on VIST. In this paper, we present…