← Search

Daniel Rubenstein

3 accepted papers

2025

BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning

NeurIPS 2025spotlight

Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-…

Cited by 0SourcecodeScholar
2025

Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis

CVPR 2025poster

We present a simple approach to make pre-trained Vision Transformers (ViTs) interpretable for fine-grained analysis, aiming to identify and localize the traits that distinguish visually similar categories, such as bird species. Pre-trained ViTs, such as DINO, have demonstrated remarkable capabilitie…

2024

A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis

ICLR 2024poster

We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an im…