← Search

Akshay Chaudhari

8 accepted papers

2026

Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)

CVPR 2026

Diffusion Magnetic Resonance Imaging (dMRI) plays a critical role in studying microstructural changes in the brain. It is, therefore, widely used in clinical practice; yet progress in learning general-purpose representations from dMRI has been limited. A key challenge is that existing deep learning

Cited by 0SourcecodeScholar
2026

Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision

CVPR 2026

Foundation models have transformed vision and language by learning general-purpose representations from large-scale unlabeled data, yet 3D medical imaging lacks analogous approaches. Existing self-supervised methods rely on low-level reconstruction or contrastive objectives that fail to capture the

Cited by 0SourcecodeScholar
2026

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

ICML 2026poster

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely assoc…

Cited by 0SourceScholar
2023

DDM$^2$: Self-Supervised Diffusion MRI Denoising with Generative Diffusion Models

ICLR 2023poster

Magnetic resonance imaging (MRI) is a common and life-saving medical imaging technique. However, acquiring high signal-to-noise ratio MRI scans requires long scan times, resulting in increased costs and patient discomfort, and decreased throughput. Thus, there is great interest in denoising MRI scan…

2023

Efficient Diagnosis Assignment Using Unstructured Clinical Notes

ACL 2023short

Electronic phenotyping entails using electronic health records (EHRs) to identify patients with specific health outcomes and determine when those outcomes occurred. Unstructured clinical notes, which contain a vast amount of information, are a valuable resource for electronic phenotyping. However, t…

Cited by 2SourcePDFScholar
2023

ViLLA: Fine-Grained Vision-Language Representation Learning from Real-World Data

ICCV 2023poster

Vision-language models (VLMs), such as CLIP and ALIGN, are generally trained on datasets consisting of image-caption pairs obtained from the web. However, real-world multimodal datasets, such as healthcare data, are significantly more complex: each image (e.g. X-ray) is often paired with text (e.g.…

Cited by 10PDFcodeScholar
2021

Designing Counterfactual Generators using Deep Model Inversion

NeurIPS 2021poster

Explanation techniques that synthesize small, interpretable changes to a given image while producing desired changes in the model prediction have become popular for introspecting black-box models. Commonly referred to as counterfactuals, the synthesized explanations are required to contain discernib…

Cited by 27SourcePDFScholar
2021

SKM-TEA: A Dataset for Accelerated MRI Reconstruction with Dense Image Labels for Quantitative Clinical Evaluation

NeurIPS 2021poster

Magnetic resonance imaging (MRI) is a cornerstone of modern medical imaging. However, long image acquisition times, the need for qualitative expert analysis, and the lack of (and difficulty extracting) quantitative indicators that are sensitive to tissue health have curtailed widespread clinical and…

Cited by 72SourcecodeScholar