← Search

Olaf Booij

5 accepted papers

2023

Objects Do Not Disappear: Video Object Detection by Single-Frame Object Location Anticipation

ICCV 2023poster

Objects in videos are typically characterized by continuous smooth motion. We exploit continuous smooth motion in three ways. 1) Improved accuracy by using object motion as an additional source of supervision, which we obtain by anticipating object locations from a static keyframe. 2) Improved effic…

Cited by 5PDFcodeScholar
2022

BoxeR: Box-Attention for 2D and 3D Transformers

CVPR 2022poster

In this paper, we propose a simple attention mechanism, we call Box-Attention. It enables spatial interaction between grid features, as sampled from boxes of interest, and improves the learning capability of transformers for several vision tasks. Specifically, we present BoxeR, short for Box Transfo…

Cited by 43PDFcodeScholar
2022

Visual Cross-View Metric Localization with Dense Uncertainty Estimates

ECCV 2022poster

"This work addresses visual cross-view metric localization for outdoor robotics. Given a ground-level color image and a satellite patch that contains the local surroundings, the task is to identify the location of the ground camera within the satellite patch. Related work addressed this task for ran…

2021

Cross-View Matching for Vehicle Localization by Learning Geographically Local Representations

RA-L 2021

Cross-view matching aims to learn a shared image representation between ground-level images and satellite or aerial images at the same locations. In robotic vehicles, matching a camera image to a database of geo-referenced aerial imagery can serve as a method for self-localization. However, existing

Cited by 27SourceScholar
2021

No Frame Left Behind: Full Video Action Recognition

CVPR 2021poster

Not all video frames are equally informative for recognizing an action. It is computationally infeasible to train deep networks on all video frames when actions develop over hundreds of frames. A common heuristic is uniformly sampling a small number of video frames and using these to recognize the a…

Cited by 60PDFcodeScholar