← Search

David J. Crandall

18 accepted papers

2025

HVPUNet: Hybrid-Voxel Point-cloud Upsampling Network

ICCV 2025poster

Point-cloud upsampling aims to generate dense point sets from sparse or incomplete 3D data. Most existing work uses a point-to-point framework. While this method achieves high geometric precision, it is slow because of irregular memory accesses to process unstructured point data. Alternatively, voxe…

Cited by 0SourcePDFScholar
2023

Correct for Whom? Subjectivity and the Evaluation of Personalized Image Aesthetics Assessment Models

AAAI 2023technical

The problem of image aesthetic quality assessment is surprisingly difficult to define precisely. Most early work attempted to estimate the average aesthetic rating of a group of observers, while some recent work has shifted to an approach based on few-shot personalization. In this paper, we connect…

Cited by 5SourcePDFScholar
2023

Polyline Generative Navigable Space Segmentation for Autonomous Visual Navigation

RA-L 2023

Detecting navigable space is a fundamental capability for mobile robots navigating in unknown or unmapped environments. In this work, we treat visual navigable space segmentation as a scene decomposition problem and propose <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www

Cited by 6SourceScholar
2020

Dynamic Dual-Attentive Aggregation Learning for Visible-Infrared Person Re-Identification

ECCV 2020poster

Visible-infrared person re-identification (VI-ReID) is a challenging cross-modality pedestrian retrieval problem. Due to the large intra-class variations and cross-modality discrepancy with large amount of sample noise, it is difficult to learn discriminative part features. Existing VI-ReID methods…

2020

HOPE-Net: A Graph-Based Model for Hand-Object Pose Estimation

CVPR 2020poster

Hand-object pose estimation (HOPE) aims to jointly detect the poses of both a hand and of a held object. In this paper, we propose a lightweight model called HOPE-Net which jointly estimates hand and object pose in 2D and 3D in real-time. Our network uses a cascade of two adaptive graph convolutiona…

Cited by 263PDFcodeScholar
2020

Learning Video Object Segmentation From Unlabeled Videos

CVPR 2020poster

We propose a new method for video object segmentation (VOS) that addresses object pattern learning from unlabeled videos, unlike most existing methods which rely heavily on extensive annotated data. We introduce a unified unsupervised/weakly supervised learning framework, called MuG, that comprehens…

Cited by 192PDFcodeScholar
2019

Automatic Annotation for Semantic Segmentation in Indoor Scenes

IROS 2019poster

Domestic robots could eventually transform our lives, but safely operating in home environments requires a rich understanding of indoor scenes. Learning-based techniques for scene segmentation require large-scale, pixel-level annotations, which are laborious and expensive to collect. We propose an a…

Cited by 8SourceScholar
2019

Egocentric Vision-based Future Vehicle Localization for Intelligent Driving Assistance Systems

ICRA 2019poster

Predicting the future location of vehicles is essential for safety-critical applications such as advanced driver assistance systems (ADAS) and autonomous driving. This paper introduces a novel approach to simultaneously predict both the location and scale of target vehicles in the first-person (egoc…

Cited by 172SourceScholar
2019

Embodied Amodal Recognition: Learning to Move to Perceive Objects

ICCV 2019poster

Passive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment and actively control the viewing angle to better understand object shapes and semantics. In this…

Cited by 75PDFScholar
2019

Temporal Recurrent Networks for Online Action Detection

ICCV 2019poster

Most work on temporal action detection is formulated as an offline problem, in which the start and end times of actions are determined after the entire video is fully observed. However, important real-time applications including surveillance and driver assistance systems require identifying actions…

Cited by 233PDFcodeScholar
2019

Unsupervised Traffic Accident Detection in First-Person Videos

IROS 2019poster

Recognizing abnormal events such as traffic violations and accidents in natural driving scenes is essential for successful autonomous driving and advanced driver assistance systems. However, most work on video anomaly detection suffers from two crucial drawbacks. First, they assume cameras are fixed…

Cited by 209SourcecodeScholar
2019

Zero-Shot Video Object Segmentation via Attentive Graph Neural Networks

ICCV 2019oral

This work proposes a novel attentive graph neural network (AGNN) for zero-shot video object segmentation (ZVOS). The suggested AGNN recasts this task as a process of iterative information fusion over video graphs. Specifically, AGNN builds a fully connected graph to efficiently represent frames as n…

Cited by 353PDFcodeScholar
2018

Joint Person Segmentation and Identification in Synchronized First- and Third-person Videos

ECCV 2018poster

In a world of pervasive cameras, public spaces are often captured from multiple perspectives by cameras of different types, both fixed and mobile. An important problem is to organize these heterogeneous collections of videos by finding connections between them, such as identifying correspondences be…

Cited by 46SourcePDFScholar
2017

Identifying First-Person Camera Wearers in Third-Person Videos

CVPR 2017poster

We consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establ…

Cited by 77PDFScholar
2015

Lending A Hand: Detecting Hands and Recognizing Activities in Complex Egocentric Interactions

ICCV 2015poster

Hands appear very often in egocentric video, and their appearance and pose give important cues about what people are doing and what they are paying attention to. But existing work in hand detection has made strong assumptions that work well in only simple scenarios, such as with limited interaction…

Cited by 541PDFScholar