← Search

Aljoša Ošep

16 accepted papers

2023

Lidar Panoptic Segmentation and Tracking without Bells and Whistles

IROS 2023poster

State-of-the-art lidar panoptic segmentation (LPS) methods follow “bottom-up” segmentation-centric fashion wherein they build upon semantic segmentation networks by utilizing clustering to obtain object instances. In this paper, we re-think this approach and propose a surprisingly simple yet effecti…

Cited by 8SourcecodeScholar
2023

Pix2map: Cross-Modal Retrieval for Inferring Street Maps From Images

CVPR 2023poster

Self-driving vehicles rely on urban street maps for autonomous navigation. In this paper, we introduce Pix2Map, a method for inferring urban street map topology directly from ego-view images, as needed to continually update and expand existing maps. This is a challenging task, as we need to infer a…

Cited by 10SourcePDFScholar
2022

DirectTracker: 3D Multi-Object Tracking Using Direct Image Alignment and Photometric Bundle Adjustment

IROS 2022poster

Direct methods have shown excellent performance in the applications of visual odometry and SLAM. In this work we propose to leverage their effectiveness for the task of 3D multi-object tracking. To this end, we propose DirectTracker, a framework that effectively combines direct image alignment for t…

Cited by 5SourceScholar
2022

Forecasting From LiDAR via Future Object Detection

CVPR 2022poster

Object detection and forecasting are fundamental components of embodied perception. These two problems, however, are largely studied in isolation by the community. In this paper, we propose an end-to-end approach for motion forecasting based on raw sensor measurement as opposed to ground truth track…

Cited by 40PDFcodeScholar
2022

Is Geometry Enough for Matching in Visual Localization?

ECCV 2022poster

"In this paper, we propose to go beyond the well-established approach to vision-based localization that relies on visual descriptor matching between a query image and a 3D point cloud. While matching keypoints via visual descriptors makes localization highly accurate, it has significant storage dema…

2022

PolarMOT: How Far Can Geometric Relations Take Us in 3D Multi-Object Tracking?

ECCV 2022poster

"Most (3D) multi-object tracking methods rely on appearance-based cues for data association. By contrast, we investigate how far we can get by only encoding geometric relationships between objects in 3D space as cues for data-driven data association. We encode 3D detections as nodes in a graph, wher…

Cited by 57SourcePDFScholar
2021

(Just) A Spoonful of Refinements Helps the Registration Error Go Down

ICCV 2021poster

In this paper, we tackle data-driven 3D point cloud registration. Given point correspondences, the standard Kabsch algorithm provides an optimal rotation estimate. This allows to train registration models in an end-to-end manner by differentiating the SVD operation. However, given the initial rotati…

Cited by 3PDFcodeScholar
2021

MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?

ICCV 2021poster

Deep learning-based methods for video pedestrian detection and tracking require large volumes of training data to achieve good performance. However, data acquisition in crowded public environments raises data privacy concerns -- we are not allowed to simply record and store data without the explicit…

Cited by 159PDFScholar
2019

Large-Scale Object Mining for Object Discovery from Unlabeled Video

ICRA 2019poster

This paper addresses the problem of object discovery from unlabeled driving videos captured in a realistic automotive setting. Identifying recurring object categories in such raw video streams is a very challenging problem. Not only do object candidates first have to be localized in the input images…

Cited by 32SourceScholar
2018

Track, Then Decide: Category-Agnostic Vision-Based Multi-Object Tracking

ICRA 2018poster

The most common paradigm for vision-based multi-object tracking is tracking-by-detection, due to the availability of reliable detectors for several important object categories such as cars and pedestrians. However, future mobile systems will need a capability to cope with rich human-made environment…

Cited by 83SourceScholar
2016

Scene flow propagation for semantic mapping and object discovery in dynamic street scenes

IROS 2016poster

Scene understanding is an important prerequisite for vehicles and robots that operate autonomously in dynamic urban street scenes. For navigation and high-level behavior planning, the robots not only require a persistent 3D model of the static surroundings—equally important, they need to perceive an…

Cited by 62SourceScholar