← Search

Leo Anthony Celi

7 accepted papers

2025

Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs

AAAI 2025technical

Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising solution to automate various procedures such as triaging, preliminary tests like visual acuity assessment, and report su…

Cited by 0SourcePDFScholar
2025

WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation

NAACL 2025findings

Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets derived from national medical examinations have long served as…

2024

Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias

NeurIPS 2024poster

Large language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in their training data. In this study, we introduce \textbf{Cross-Care}, the first benchmark framework dedicated to assessin…

2024

Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks

EMNLP 2024finding

Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To st…

2024

MedDec: A Dataset for Extracting Medical Decisions from Discharge Summaries

ACL 2024findings

Medical decisions directly impact individuals’ health and well-being. Extracting decision spans from clinical notes plays a crucial role in understanding medical decision-making processes. In this paper, we develop a new dataset called “MedDec,” which contains clinical notes of eleven different phen…

2021

Chest ImaGenome Dataset for Clinical Reasoning

NeurIPS 2021poster

Despite the progress in automatic detection of radiologic findings from Chest X-Ray (CXR) images in recent years, a quantitative evaluation of the explainability of these models is hampered by the lack of locally labeled datasets for different findings. With the exception of a few expert-labeled sma…

Cited by 76SourcecodeScholar
2020

Expert-Supervised Reinforcement Learning for Offline Policy Learning and Evaluation

NeurIPS 2020poster

Offline Reinforcement Learning (RL) is a promising approach for learning optimal policies in environments where direct exploration is expensive or unfeasible. However, the adoption of such policies in practice is often challenging, as they are hard to interpret within the application context, and la…