← Search

Yunsoo Kim

6 accepted papers

2026

Error Correction in Radiology Reports: A Knowledge Distillation-Based Multi-Stage Framework

AAAI 2026technical

The increasing complexity and workload of clinical radiology leads to inevitable oversights and mistakes in their use as diagnostic tools, causing delayed treatments and sometimes life-threatening harm to patients. While large language models (LLMs) have shown remarkable progress in many tasks, thei

Cited by 0SourcePDFScholar
2025

BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain

ACL 2025finding

Biomedical reasoning often requires traversing interconnected relationships across entities such as drugs, diseases, and proteins. Despite the increasing prominence of large language models (LLMs), existing benchmarks lack the ability to evaluate multi-hop reasoning in the biomedical domain, particu…

Cited by 0SourcePDFScholar
2025

HARE: an entity and relation centric evaluation framework for histopathology reports

EMNLP 2025

Medical domain automated text generation is an active area of research and development; however, evaluating the clinical quality of generated reports remains a challenge, especially in instances where domain-specific metrics are lacking, e.g. histopathology. We propose HARE (Histopathology Automated

2025

Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

ACL 2025finding

Recent advancements in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, particularly in generating radiology reports from chest X-rays (CXR). However, these models still suffer from hallucinations and clinically significant errors, limitin…

Cited by 0SourcePDFScholar
2023

Chemical Language Understanding Benchmark

ACL 2023industry

In this paper, we introduce the benchmark datasets named CLUB (Chemical Language Understanding Benchmark) to facilitate NLP research in the chemical industry. We have 4 datasets consisted of text and token classification tasks. As far as we have recognized, it is one of the first examples of chemica…

Cited by 4SourcePDFScholar