← Search

Wenjia Bai

7 accepted papers

2026

Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases

ICML 2026poster

Clinical abnormality grounding for rare diseases is often hindered by data scarcity, rendering supervised fine-tuning infeasible and single-pass inference highly unstable. Thus, we propose Dynamic Decision Learning (DDL), a framework that enables frozen LVLMs to refine their decisions across languag…

Cited by 0SourceScholar
2026

Virtual Full-stack Scanning of Brain MRI via Imputing Any Quantised Code

CVPR 2026

Magnetic resonance imaging (MRI) is a powerful and versatile imaging technique, offering a wide spectrum of information about the anatomy by employing different acquisition modalities. However, in the clinical workflow, it is impractical to collect all relevant modalities due to the scan time and co

Cited by 0SourcecodeScholar
2025

Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputs

EMNLP 2025

Recent advancements in multimodal representation learning for electrocardiogram (ECG) have moved onto learning representations by aligning ECG signals with their paired free-text reports. However, current methods often result in suboptimal alignment of ECG signals with their corresponding text repor

2025

NOVA: A Benchmark for Rare Anomaly Localization and Clinical Reasoning in Brain MRI

NeurIPS 2025oral

In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Open-world recognition ensures that such systems remain robust as ever-emerging, previously _unknown_ categories appear and must be addressed without retraining. Foundation and vision-la…

Cited by 0SourceScholar
2024

G2D: From Global to Dense Radiography Representation Learning via Vision-Language Pre-training

NeurIPS 2024poster

Medical imaging tasks require an understanding of subtle and localized visual features due to the inherently detailed and area-specific nature of pathological patterns, which are crucial for clinical diagnosis. Although recent advances in medical vision-language pre-training (VLP) enable models to l…

2024

TIP: Tabular-Image Pre-training for Multimodal Classification with Incomplete Data

ECCV 2024poster

"Images and structured tables are essential parts of real-world databases. Though tabular-image representation learning is promising for creating new insights, it remains a challenging task, as tabular data is typically heterogeneous and incomplete, presenting significant modality disparities with i…

2024

Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge Enhancement

ICML 2024poster

Electrocardiograms (ECGs) are non-invasive diagnostic tools crucial for detecting cardiac arrhythmic diseases in clinical practice. While ECG Self-supervised Learning (eSSL) methods show promise in representation learning from unannotated ECG data, they often overlook the clinical knowledge that can…