← Search

Stephanie Fu

8 accepted papers

2026

Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing

CVPR 2026

Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos---they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant spatiotemporal redundancy. We introduce AutoGaze, a lightweigh

Cited by 0SourcecodeScholar
2024

A Vision Check-up for Language Models

CVPR 2024highlight

What does learning to model relationships between strings teach Large Language Models (LLMs) about the visual world? We systematically evaluate LLMs' abilities to generate and recognize an assortment of visual concepts of increasing complexity and then demonstrate how a preliminary visual representa…

Cited by 29SourcePDFScholar
2024

Evaluating Multiview Object Consistency in Humans and Image Models

NeurIPS 2024poster

We introduce a benchmark to directly evaluate the alignment between human observers and vision models on a 3D shape inference task. We leverage an experimental design from the cognitive sciences: given a set of images, participants identify which contain the same/different objects, despite considera…

Cited by 5SourcecodeScholar
2024

FeatUp: A Model-Agnostic Framework for Features at Any Resolution

ICLR 2024poster

Deep features are a cornerstone of computer vision research, capturing image semantics and enabling the community to solve downstream tasks even in the zero- or few-shot regime. However, these features often lack the spatial resolution to directly perform dense prediction tasks like segmentation and…

2024

OpenStreetView-5M: The Many Roads to Global Visual Geolocation

CVPR 2024poster

Determining the location of an image anywhere on Earth is a complex visual task which makes it particularly relevant for evaluating computer vision algorithms. Determining the location of an image anywhere on Earth is a complex visual task which makes it particularly relevant for evaluating computer…

2024

When does perceptual alignment benefit vision representations?

NeurIPS 2024poster

Humans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range of semantic abstractions but improperly weigh these attributes and thus make inferences misaligned with human perceptio…

Cited by 5SourcePDFScholar
2023

DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

NeurIPS 2023spotlight

Current perceptual similarity metrics operate at the level of pixels and patches. These metrics compare images in terms of their low-level colors and textures, but fail to capture mid-level similarities and differences in image layout, object pose, and semantic content. In this paper, we develop a p…

2022

Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning

ICLR 2022poster

Visual search, recommendation, and contrastive similarity learning power technologies that impact billions of users worldwide. Modern model architectures can be complex and difficult to interpret, and there are several competing techniques one can use to explain a search engine's behavior. We show t…

Cited by 9SourcePDFScholar