← Search

Richard J. Chen

12 accepted papers

2026

Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance Learning

ICLR 2026poster

Multiple Instance Learning (MIL) is the predominant approach for classifying gigapixel whole-slide images in computational pathology. MIL follows a sequence of 1) extracting patch features, 2) applying a linear layer to obtain task-specific patch features, and 3) aggregating the patches into a slide…

Cited by 0SourcecodeScholar
2025

Do Multiple Instance Learning Models Transfer?

ICML 2025spotlight

Multiple Instance Learning (MIL) is a cornerstone approach in computational pathology for distilling embeddings from gigapixel tissue images into patient-level representations to predict clinical outcomes. However, MIL is frequently challenged by the constraints of working with small, weakly-supervi…

Cited by 0SourcePDFScholar
2024

HEST-1k: A Dataset For Spatial Transcriptomics and Histology Image Analysis

NeurIPS 2024spotlight

Spatial transcriptomics enables interrogating the molecular composition of tissue with ever-increasing resolution and sensitivity. However, costs, rapidly evolving technology, and lack of standards have constrained computational methods in ST to narrow tasks and small cohorts. In addition, the under…

2024

Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival Prediction

CVPR 2024poster

Integrating whole-slide images (WSIs) and bulk transcriptomics for predicting patient survival can improve our understanding of patient prognosis. However this multimodal task is particularly challenging due to the different nature of these data: WSIs represent a very high-dimensional spatial descri…

2024

Morphological Prototyping for Unsupervised Slide Representation Learning in Computational Pathology

CVPR 2024poster

Representation learning of pathology whole-slide images (WSIs) has been has primarily relied on weak supervision with Multiple Instance Learning (MIL). However the slide representations resulting from this approach are highly tailored to specific clinical tasks which limits their expressivity and ge…

2024

Multimodal Prototyping for cancer survival prediction

ICML 2024poster

Multimodal survival methods combining gigapixel histology whole-slide images (WSIs) and transcriptomic profiles are particularly promising for patient prognostication and stratification. Current approaches involve tokenizing the WSIs into smaller patches ($>10^4$ patches) and transcriptomics into ge…

2024

Multistain Pretraining for Slide Representation Learning in Pathology

ECCV 2024poster

"Developing self-supervised learning (SSL) models that can learn universal and transferable representations of H&E gigapixel whole-slide images (WSIs) is becoming increasingly valuable in computational pathology. These models hold the potential to advance critical tasks such as few-shot classificati…

2024

Transcriptomics-guided Slide Representation Learning in Computational Pathology

CVPR 2024poster

Self-supervised learning (SSL) has been successful in building patch embeddings of small histology images (e.g. 224 x 224 pixels) but scaling these models to learn slide embeddings from the entirety of giga-pixel whole-slide images (WSIs) remains challenging. Here we leverage complementary informati…

2023

Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework

NeurIPS 2023poster

The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain fundamental research questions: How can we quantify the interact…

2023

Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology Images

CVPR 2023poster

Contrastive visual language pretraining has emerged as a powerful method for either training new language-aware image encoders or augmenting existing pretrained models with zero-shot visual recognition capabilities. However, existing works typically train on large datasets of image-text pairs and ha…

2022

Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning

CVPR 2022oral

Vision Transformers (ViTs) and their multi-scale and hierarchical variations have been successful at capturing image representations but their use has been generally studied for low-resolution images (e.g. - 256x256, 384x384). For gigapixel whole-slide imaging (WSI) in computational pathology, WSIs…

Cited by 556PDFcodeScholar
2021

Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide Images

ICCV 2021poster

Survival outcome prediction is a challenging weakly-supervised and ordinal regression task in computational pathology that involves modeling complex interactions within the tumor microenvironment in gigapixel whole slide images (WSIs). Despite recent progress in formulating WSIs as bags for multiple…

Cited by 301PDFcodeScholar