← Search

Drew A. Hudson

9 accepted papers

2026

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

CVPR 2026

Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features are poorly aligned across different modalities. For instance, the feature embedding for an RGB image and its corresponding depth map of the same scene exhibit

Cited by 0SourcecodeScholar
2025

LayerLock: Non-collapsing Representation Learning with Progressive Freezing

ICCV 2025poster

We introduce LayerLock, a simple yet effective approach for self-supervised visual representation learning, that gradually transitions throughout training from predicting shallow features to deeper ones through progressive layer freezing. First, we make the observation that during training of video…

Cited by 0SourcePDFScholar
2024

Moving Off-the-Grid: Scene-Grounded Video Representations

NeurIPS 2024spotlight

Current vision models typically maintain a fixed correspondence between their representation structure and image space. Each layer comprises a set of tokens arranged “on-the-grid,” which biases patches or tokens to encode information at a specific spatio(-temporal) location. In this work we present…

Cited by 2SourcePDFScholar
2024

Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models

NeurIPS 2024spotlight

We address the problem of multi-object 3D pose control in image diffusion models. Instead of conditioning on a sequence of text tokens, we propose to use a set of per-object representations, *Neural Assets*, to control the 3D pose of individual objects in a scene. Neural Assets are obtained by pooli…

Cited by 13SourcePDFScholar
2024

SODA: Bottleneck Diffusion Models for Representation Learning

CVPR 2024poster

We introduce SODA a self-supervised diffusion model designed for representation learning. The model incorporates an image encoder which distills a source view into a compact representation that in turn guides the generation of related novel views. We show that by imposing a tight bottleneck between…

2019

GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

CVPR 2019oral

We introduce GQA, a new dataset for real-world visual reasoning and compositional question answering, seeking to address key shortcomings of previous VQA datasets. We have developed a strong and robust question engine that leverages Visual Genome scene graph structures to create 22M diverse reasonin…

Cited by 2286PDFcodeScholar