← Search

Sara Beery

18 accepted papers

2026

A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn’t)

ICML 2026poster

Instruction fine-tuning of large language models (LLMs) often involves selecting a subset of instruction training data from a large candidate pool, using a small query set from the target task. Despite growing interest, the literature on targeted instruction selection remains fragmented and opaque: …

Cited by 0SourceScholar
2026

PRISM: Controllable Diffusion for Compound Image Restoration with Scientific Fidelity

ICLR 2026poster

Scientific and environmental imagery are often degraded by multiple compounding factors related to sensor noise and environmental effects. Existing restoration methods typically treat these mixed effects by iteratively removing fixed categories, lacking the compositionality needed to handle real-wor…

Cited by 0SourceScholar
2025

Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations

NeurIPS 2025spotlight

Benchmarks for out-of-distribution (OOD) generalization often reveal a strong positive correlation between in-distribution (ID) and OOD accuracy across models, a phenomenon known as “accuracy-on-the-line.” This pattern is commonly interpreted as evidence that spurious correlations—relationships that…

Cited by 0SourceScholar
2025

Consensus-Driven Active Model Selection

ICCV 2025poster

The widespread availability of off-the-shelf machine learning models poses a challenge: which model, of the many available candidates, should be chosen for a given data analysis task? This question of model selection is traditionally answered by collecting and annotating a validation dataset---a cos…

2025

Open-Insect: Benchmarking Open-Set Recognition of Novel Species in Biodiversity Monitoring

NeurIPS 2025spotlight

Global biodiversity is declining at an unprecedented rate, yet little information is known about most species and how their populations are changing. Indeed, some 90% Earth’s species are estimated to be completely unknown. Machine learning has recently emerged as a promising tool to facilitate long-…

Cited by 0SourceScholar
2025

Personalized Representation from Personalized Generation

ICLR 2025poster

Modern vision models excel at general purpose downstream tasks. It is unclear, however, how they may be used for personalized vision tasks, which are both fine-grained and data-scarce. Recent works have successfully applied synthetic data to general-purpose representation learning, while advances in…

2025

Visually Consistent Hierarchical Image Classification

ICLR 2025poster

Hierarchical classification predicts labels across multiple levels of a taxonomy, e.g., from coarse-level \textit{Bird} to mid-level \textit{Hummingbird} to fine-level \textit{Green hermit}, allowing flexible recognition under varying visual conditions. It is commonly framed as multiple single-leve…

Cited by 0SourcePDFScholar
2024

Are They the Same Picture? Adapting Concept Bottleneck Models for Human-AI Collaboration in Image Retrieval

IJCAI 2024poster

Image retrieval plays a pivotal role in applications from wildlife conservation to healthcare, for finding individual animals or relevant images to aid diagnosis. Although deep learning techniques for image retrieval have advanced significantly, their imperfect real-world performance often necessita…

2024

INQUIRE: A Natural World Text-to-Image Retrieval Benchmark

NeurIPS 2024poster

We introduce INQUIRE, a text-to-image retrieval benchmark designed to challenge multimodal vision-language models on expert-level queries. INQUIRE includes iNaturalist 2024 (iNat24), a new dataset of five million natural world images, along with 250 expert-level retrieval queries. These queries are…

2024

Position: Application-Driven Innovation in Machine Learning

ICML 2024poster

In this position paper, we argue that application-driven research has been systemically under-valued in the machine learning community. As applications of machine learning proliferate, innovative algorithms inspired by specific real-world challenges have become increasingly important. Such work offe…

Cited by 4SourcePDFScholar
2023

MammalNet: A Large-Scale Video Benchmark for Mammal Recognition and Behavior Understanding

CVPR 2023poster

Monitoring animal behavior can facilitate conservation efforts by providing key insights into wildlife health, population status, and ecosystem function. Automatic recognition of animals and their behaviors is critical for capitalizing on the large unlabeled datasets generated by modern video device…

2022

Extending the WILDS Benchmark for Unsupervised Adaptation

ICLR 2022oral

Machine learning systems deployed in the wild are often trained on a source distribution but deployed on a different target distribution. Unlabeled data can be a powerful point of leverage for mitigating these distribution shifts, as it is frequently much more available than labeled data and can oft…

Cited by 143SourcePDFScholar
2022

The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain Shift

CVPR 2022poster

Generalization to novel domains is a fundamental challenge for computer vision. Near-perfect accuracy on benchmarks is common, but these models do not work as expected when deployed outside of the training distribution. To build computer vision systems that truly solve real-world problems at global…

Cited by 53PDFScholar
2022

The Caltech Fish Counting Dataset: A Benchmark for Multiple-Object Tracking and Counting

ECCV 2022poster

"We present the Caltech Fish Counting Dataset (CFC), a large-scale dataset for detecting, tracking, and counting fish in sonar videos. We identify sonar videos as a rich source of data for advancing low signal-to-noise computer vision applications and tackling domain generalization in multiple-objec…

2021

Benchmarking Representation Learning for Natural World Image Collections

CVPR 2021poster

Recent progress in self-supervised learning has resulted in models that are capable of extracting rich representations from image collections without requiring any explicit label supervision. However, to date the vast majority of these approaches have restricted themselves to training on standard be…

Cited by 191PDFcodeScholar
2020

Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection

CVPR 2020poster

In static monitoring cameras, useful contextual information can stretch far beyond the few seconds typical video understanding models might see: subjects may exhibit similar behavior over multiple days, and background objects remain static. Due to power and storage constraints, sampling frequencies…

Cited by 164PDFScholar