← Search

Grant Van Horn

23 accepted papers

2026

RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs

CVPR 2026

Fine-grained bird species identification in the wild is frequently unanswerable from a single image: key cues may be non-visual (e.g. vocalization), or obscured due to occlusion, camera angle, or low resolution. Yet today's multimodal systems are typically judged on answerable, in-schema cases, enco

Cited by 0SourcecodeScholar
2025

CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing

NeurIPS 2025poster

Mastering fine-grained visual recognition, essential in many expert domains, can require that specialists undergo years of dedicated training. Modeling the progression of such expertize in humans remains challenging, and accurately inferring a human learner’s knowledge state is a key step toward und…

Cited by 0SourceScholar
2025

Consensus-Driven Active Model Selection

ICCV 2025poster

The widespread availability of off-the-shelf machine learning models poses a challenge: which model, of the many available candidates, should be chosen for a given data analysis task? This question of model selection is traditionally answered by collecting and annotating a validation dataset---a cos…

2025

Feedforward Few-shot Species Range Estimation

ICML 2025poster

Knowing where a particular species can or cannot be found on Earth is crucial for ecological research and conservation efforts. By mapping the spatial ranges of all species, we would obtain deeper insights into how global biodiversity is affected by climate change and habitat loss. However, accurat…

Cited by 0SourcePDFScholar
2025

Generate, Transduct, Adapt: Iterative Transduction with VLMs

ICCV 2025poster

Transductive zero-shot learning with vision-language models leverages image-image similarities within the dataset to achieve better classification accuracy compared to the inductive setting. However, there is little work that explores the structure of the language space in this context. We propose G…

Cited by 0SourcePDFScholar
2025

WildSAT: Learning Satellite Image Representations from Wildlife Observations

ICCV 2025poster

Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with millions of geo-tagged wildlife observations readily-available on c…

2024

Combining Observational Data and Language for Species Range Estimation

NeurIPS 2024poster

Species range maps (SRMs) are essential tools for research and policy-making in ecology, conservation, and environmental management. However, traditional SRMs rely on the availability of environmental covariates and high-quality observational data, both of which can be challenging to obtain due to g…

2024

Human-in-the-Loop Visual Re-ID for Population Size Estimation

ECCV 2024poster

"Computer vision-based re-identification (Re-ID) systems are increasingly being deployed for estimating population size in large image collections. However, the estimated size can be significantly inaccurate when the task is challenging or when deployed on data from new distributions. We propose a h…

2024

INQUIRE: A Natural World Text-to-Image Retrieval Benchmark

NeurIPS 2024poster

We introduce INQUIRE, a text-to-image retrieval benchmark designed to challenge multimodal vision-language models on expert-level queries. INQUIRE includes iNaturalist 2024 (iNat24), a new dataset of five million natural world images, along with 250 expert-level retrieval queries. These queries are…

2024

Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions

CVPR 2024poster

The zero-shot performance of existing vision-language models (VLMs) such as CLIP is limited by the availability of large-scale aligned image and text datasets in specific domains. In this work we leverage two complementary sources of information -- descriptions of categories generated by large langu…

2023

Active Learning-Based Species Range Estimation

NeurIPS 2023poster

We propose a new active learning approach for efficiently estimating the geographic range of a species from a limited number of on the ground observations. We model the range of an unmapped species of interest as the weighted combination of estimated ranges obtained from a set of different species.…

2023

Spatial Implicit Neural Representations for Global-Scale Species Mapping

ICML 2023poster

Estimating the geographical range of a species from sparse observations is a challenging and important geospatial prediction problem. Given a set of locations where a species has been observed, the goal is to build a model to predict whether the species is present or absent at any location. This pro…

2022

Exploring Fine-Grained Audiovisual Categorization with the SSW60 Dataset

ECCV 2022poster

"We present a new benchmark dataset, Sapsucker Woods 60 (SSW60), for advancing research on audiovisual fine-grained categorization. While our community has made great strides in fine-grained visual categorization on images, the counterparts in audio and video fine-grained categorization are relative…

2022

On Label Granularity and Object Localization

ECCV 2022poster

"Weakly supervised object localization (WSOL) aims to learn representations that encode object location using only image-level category labels. However, many objects can be labeled at different levels of granularity. Is it an animal, a bird, or a great horned owl? Which image-level labels should we…

2022

The Caltech Fish Counting Dataset: A Benchmark for Multiple-Object Tracking and Counting

ECCV 2022poster

"We present the Caltech Fish Counting Dataset (CFC), a large-scale dataset for detecting, tracking, and counting fish in sonar videos. We identify sonar videos as a rich source of data for advancing low signal-to-noise computer vision applications and tackling domain generalization in multiple-objec…

2021

Benchmarking Representation Learning for Natural World Image Collections

CVPR 2021poster

Recent progress in self-supervised learning has resulted in models that are capable of extracting rich representations from image collections without requiring any explicit label supervision. However, to date the vast majority of these approaches have restricted themselves to training on standard be…

Cited by 191PDFcodeScholar
2018

The INaturalist Species Classification and Detection Dataset

CVPR 2018poster

Existing image classification datasets used in computer vision tend to have a uniform distribution of images across object categories. In contrast, the natural world is heavily imbalanced, as some species are more abundant and easier to photograph than others. To encourage further progress in challe…

2015

Building a Bird Recognition App and Large Scale Dataset With Citizen Scientists: The Fine Print in Fine-Grained Dataset Collection

CVPR 2015poster

We introduce tools and methodologies to collect high quality, large scale fine-grained computer vision datasets using citizen scientists -- crowd annotators who are passionate and knowledgeable about specific domains such as birds or airplanes. We worked with citizen scientists and domain experts t…

Cited by 737SourcePDFScholar