← Search

Nigam Shah

9 accepted papers

2026

Position: Benchmarks Do Not Measure Deployment Readiness in Clinical AI

ICML 2026poster

Despite large language models (LLMs) achieving impressive performance on benchmark tasks such as medical question answering, their real-world utility remains limited. We argue that while benchmarks play a valuable role in developing methods and filtering promising models during development, they oft…

Cited by 0SourceScholar
2025

Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHR Data

ICLR 2025poster

Foundation Models (FMs) trained on Electronic Health Records (EHRs) have achieved state-of-the-art results on numerous clinical prediction tasks. However, prior EHR FMs typically have context windows of $<$1k tokens, which prevents them from modeling full patient EHRs which can exceed 10k's of event…

Cited by 1SourcePDFScholar
2025

STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology

NeurIPS 2025poster

Multi-class tissue-type classification of colorectal cancer (CRC) histopathologic images is a significant step in the development of downstream machine learning models for diagnosis and treatment planning. However, publicly available CRC datasets used to build tissue classifiers often suffer from in…

Cited by 0SourceScholar
2025

Time-to-Event Pretraining for 3D Medical Imaging

ICLR 2025poster

With the rise of medical foundation models and the growing availability of imaging data, scalable pretraining techniques offer a promising way to identify imaging biomarkers predictive of future disease risk. While current self-supervised methods for 3D medical imaging models capture local structura…

2024

MOTOR: A Time-to-Event Foundation Model For Structured Medical Records

ICLR 2024spotlight

We present a self-supervised, time-to-event (TTE) foundation model called MOTOR (Many Outcome Time Oriented Representations) which is pretrained on timestamped sequences of events in electronic health records (EHR) and health insurance claims. TTE models are used for estimating the probability distr…

Cited by 17SourcePDFScholar
2024

WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks

NeurIPS 2024poster

Existing ML benchmarks lack the depth and diversity of annotations needed for evaluating models on business process management (BPM) tasks. BPM is the practice of documenting, measuring, improving, and automating enterprise workflows. However, research has focused almost exclusively on one task -- f…

Cited by 1SourcecodeScholar
2023

EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models

NeurIPS 2023spotlight

While the general machine learning (ML) community has benefited from public datasets, tasks, and models, the progress of ML in healthcare has been hampered by a lack of such shared assets. The success of foundation models creates new challenges for healthcare ML by requiring access to shared pretrai…

2023

Efficient Diagnosis Assignment Using Unstructured Clinical Notes

ACL 2023short

Electronic phenotyping entails using electronic health records (EHRs) to identify patients with specific health outcomes and determine when those outcomes occurred. Unstructured clinical notes, which contain a vast amount of information, are a valuable resource for electronic phenotyping. However, t…

Cited by 2SourcePDFScholar
2023

INSPECT: A Multimodal Dataset for Pulmonary Embolism Diagnosis and Prognosis

NeurIPS 2023poster

Synthesizing information from various data sources plays a crucial role in the practice of modern medicine. Current applications of artificial intelligence in medicine often focus on single-modality data due to a lack of publicly available, multimodal medical datasets. To address this limitation, we…

Cited by 11SourcePDFScholar