← Search

Jack Gallifant

8 accepted papers

2025

ACES: Automatic Cohort Extraction System for Event-Stream Datasets

ICLR 2025poster

Reproducibility remains a significant challenge in machine learning (ML) for healthcare. Datasets, model pipelines, and even task or cohort definitions are often private in this field, leading to a significant barrier in sharing, iterating, and understanding ML results on electronic health record (E…

2025

KScope: A Framework for Characterizing the Knowledge Status of Language Models

NeurIPS 2025poster

Characterizing a large language model's (LLM's) knowledge of a given question is challenging. As a result, prior work has primarily examined LLM behavior under knowledge conflicts, where the model's internal parametric memory contradicts information in the external context. However, this does not fu…

Cited by 0SourceScholar
2025

Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs

AAAI 2025technical

Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising solution to automate various procedures such as triaging, preliminary tests like visual acuity assessment, and report su…

Cited by 0SourcePDFScholar
2025

Sparse Autoencoder Features for Classifications and Transferability

EMNLP 2025

Sparse Autoencoders (SAEs) provide potential for uncovering structured, human-interpretable representations in Large Language Models (LLMs), making them a crucial tool for transparent and controllable AI systems. We systematically analyze SAE for interpretable feature extraction from LLMs in safety-

2025

WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation

NAACL 2025findings

Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets derived from national medical examinations have long served as…

2024

A Closer Look at AUROC and AUPRC under Class Imbalance

NeurIPS 2024poster

In machine learning (ML), a widespread claim is that the area under the precision-recall curve (AUPRC) is a superior metric for model comparison to the area under the receiver operating characteristic (AUROC) for tasks with class imbalance. This paper refutes this notion on two fronts. First, we the…

Cited by 46SourcePDFScholar
2024

Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias

NeurIPS 2024poster

Large language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in their training data. In this study, we introduce \textbf{Cross-Care}, the first benchmark framework dedicated to assessin…

2024

Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks

EMNLP 2024finding

Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To st…