← Search

David Crandall

13 accepted papers

2026

NN-kNN for Regression: Accurate Prediction from Interpretable Retrieval

IJCAI 2026

Neural Network k-Nearest Neighbor (NN-kNN) was proposed as an interpretable network model that learns feature weights and similarity to retrieve relevant cases for classification. This paper extends it to regression with the goal of generating accurate predictions based on neighboring cases with sim

Cited by 0Scholar
2025

What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning

ICCV 2025poster

Understanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or erroneous. Existing work has studied procedure-aware video representations by modeling…

2024

SePaint: Semantic Map Inpainting via Multinomial Diffusion

IROS 2024poster

Prediction beyond partial observations is crucial for robots to navigate in unknown environments because it can provide extra information regarding the surroundings beyond the current sensing range or resolution. In this work, we consider the inpainting of semantic Bird’s-Eye-View maps. We propose S…

Cited by 2SourceScholar
2023

Few-Shot Segmentation and Semantic Segmentation for Underwater Imagery

IROS 2023poster

This paper tackles image segmentation problems for underwater environments. First, we introduce a novel under-water animal-centric dataset with dense pixel-level annotations containing diverse fine-grained animal categories to mitigate the lack of diverse categories in the existing benchmarks. Then,…

Cited by 5SourcecodeScholar
2023

VindLU: A Recipe for Effective Video-and-Language Pretraining

CVPR 2023poster

The last several years have witnessed remarkable progress in video-and-language (VidL) understanding. However, most modern VidL approaches use complex and specialized model architectures and sophisticated pretraining protocols, making the reproducibility, analysis and comparisons of these frameworks…

2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2021

Error Diagnosis of Deep Monocular Depth Estimation Models

IROS 2021poster

Estimating depth from a monocular image is an ill-posed problem: when the camera projects a 3D scene onto a 2D plane, depth information is inherently and permanently lost. Nevertheless, recent work has shown impressive results in estimating 3D structure from 2D images using deep learning. In this pa…

Cited by 3SourceScholar
2020

Interaction Graphs for Object Importance Estimation in On-road Driving Videos

ICRA 2020poster

A vehicle driving along the road is surrounded by many objects, but only a small subset of them influence the driver's decisions and actions. Learning to estimate the importance of each object on the driver's real-time decision-making may help better understand human driving behavior and lead to mor…

Cited by 34SourceScholar
2019

A Self Validation Network for Object-Level Human Attention Estimation

NeurIPS 2019poster

Due to the foveated nature of the human vision system, people can focus their visual attention on a small region of their visual field at a time, which usually contains only a single object. Estimating this object of attention in first-person (egocentric) videos is useful for many human-centered rea…

2019

Meta-Reinforced Synthetic Data for One-Shot Fine-Grained Visual Recognition

NeurIPS 2019poster

This paper studies the task of one-shot fine-grained recognition, which suffers from the problem of data scarcity of novel fine-grained classes. To alleviate this problem, a off-the-shelf image generator can be applied to synthesize additional images to help one-shot learning. However, such synthesi…

2016

Stochastic Multiple Choice Learning for Training Diverse Deep Ensembles

NeurIPS 2016poster

Many practical perception systems exist within larger processes which often include interactions with users or additional components that are capable of evaluating the quality of predicted solutions. In these contexts, it is beneficial to provide these oracle mechanisms with multiple highly likely h…

Cited by 233SourcePDFScholar