← Search

Honghan Wu

7 accepted papers

2026

Error Correction in Radiology Reports: A Knowledge Distillation-Based Multi-Stage Framework

AAAI 2026technical

The increasing complexity and workload of clinical radiology leads to inevitable oversights and mistakes in their use as diagnostic tools, causing delayed treatments and sometimes life-threatening harm to patients. While large language models (LLMs) have shown remarkable progress in many tasks, thei

Cited by 0SourcePDFScholar
2025

Adverse Event Extraction from Discharge Summaries: A New Dataset, Annotation Scheme, and Initial Findings

ACL 2025long

In this work, we present a manually annotated corpus for Adverse Event (AE) extraction from discharge summaries of elderly patients, a population often underrepresented in clinical NLP resources. The dataset includes 14 clinically significant AEs—such as falls, delirium, and intracranial haemorrhage…

2025

BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain

ACL 2025finding

Biomedical reasoning often requires traversing interconnected relationships across entities such as drugs, diseases, and proteins. Despite the increasing prominence of large language models (LLMs), existing benchmarks lack the ability to evaluate multi-hop reasoning in the biomedical domain, particu…

Cited by 0SourcePDFScholar
2025

HARE: an entity and relation centric evaluation framework for histopathology reports

EMNLP 2025

Medical domain automated text generation is an active area of research and development; however, evaluating the clinical quality of generated reports remains a challenge, especially in instances where domain-specific metrics are lacking, e.g. histopathology. We propose HARE (Histopathology Automated

2025

Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

ACL 2025finding

Recent advancements in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, particularly in generating radiology reports from chest X-rays (CXR). However, these models still suffer from hallucinations and clinically significant errors, limitin…

Cited by 0SourcePDFScholar
2024

CMDL: A Large-Scale Chinese Multi-Defendant Legal Judgment Prediction Dataset

ACL 2024findings

Legal Judgment Prediction (LJP) has attracted significant attention in recent years. However, previous studies have primarily focused on cases involving only a single defendant, skipping multi-defendant cases due to complexity and difficulty. To advance research, we introduce CMDL, a large-scale rea…

2022

Quantifying Health Inequalities Induced by Data and AI Models

IJCAI 2022poster

AI technologies are being increasingly tested and applied in critical environments including healthcare. Without an effective way to detect and mitigate AI induced inequalities, AI might do more harm than good, potentially leading to the widening of underlying inequalities. This paper proposes a gen…