← Search

David Eigen

5 accepted papers

2026

Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing

CVPR 2026

Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos---they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant spatiotemporal redundancy. We introduce AutoGaze, a lightweigh

Cited by 0SourcecodeScholar
2019

Finding Task-Relevant Features for Few-Shot Learning by Category Traversal

CVPR 2019oral

Few-shot learning is an important area of research. Conceptually, humans are readily able to understand new concepts given just a few examples, while in more pragmatic terms, limited-example training situations are common practice. Recent effective approaches to few-shot learning employ a metric-le…

Cited by 477PDFcodeScholar
2015

End-to-End Integration of a Convolution Network, Deformable Parts Model and Non-Maximum Suppression

CVPR 2015poster

Deformable Parts Models and Convolutional Networks each have achieved notable performance in object detection. Yet these two approaches find their strengths in complementary areas: DPMs are well-versed in object composition, modeling fine-grained spatial relationships between parts; likewise, Conv…

Cited by 117SourcePDFScholar
2015

Predicting Depth, Surface Normals and Semantic Labels With a Common Multi-Scale Convolutional Architecture

ICCV 2015poster

In this paper we address three different computer vision tasks using a single basic architecture: depth prediction, surface normal estimation, and semantic labeling. We use a multiscale convolutional network that is able to adapt easily to each task using only small modifications, regressing from…

Cited by 3530PDFScholar
2015

Unsupervised Learning of Spatiotemporally Coherent Metrics

ICCV 2015poster

Current state-of-the-art classification and detection algorithms train deep convolutional networks using labeled data. In this work we study unsupervised feature learning with convolutional networks in the context of temporally coherent unlabeled data. We focus on feature learning from unlabeled vid…

Cited by 192PDFScholar