← Search

Caleb Chen Cao

5 accepted papers

2026

KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation

CVPR 2026

Cross-modal biomedical signals such as pathology and genomics can provide richer and more robust semantic guidance for medical image representation learning. However, the availability of such guidance remains limited, as privacy constraints and acquisition costs severely restrict access to medical i

Cited by 0SourceScholar
2026

URICA: A Uniformity Region Affine Identifier Capture Algorithm for Arbitrary Region Retrieval in Pathology Images

CVPR 2026

Whole slide image (WSI) region retrieval remains an open challenge in computational pathology, as existing methods struggle to represent and preserve information of all possible regions. Current approaches that rely on fixed-size patches or slide-level retrieval are misaligned with real clinical wor

Cited by 0SourceScholar
2023

Towards Fine-Grained Explainability for Heterogeneous Graph Neural Network

AAAI 2023technical

Heterogeneous graph neural networks (HGNs) are prominent approaches to node classification tasks on heterogeneous graphs. Despite the superior performance, insights about the predictions made from HGNs are obscure to humans. Existing explainability techniques are mainly proposed for GNNs on homogene…

2023

Two-stage holistic and contrastive explanation of image classification

UAI 2023poster

The need to explain the output of a deep neural network classifier is now widely recognized. While previous methods typically explain a single class in the output, we advocate explaining the whole output, which is a probability distribution over multiple classes. A whole-output explanation can help…

2023

ViT-CX: Causal Explanation of Vision Transformers

IJCAI 2023poster

Despite the popularity of Vision Transformers (ViTs) and eXplainable AI (XAI), only a few explanation methods have been designed specially for ViTs thus far. They mostly use attention weights of the [CLS] token on patch embeddings and often produce unsatisfactory saliency maps. This paper proposes a…