← Search

Aljosa Osep

15 accepted papers

2026

VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale

CVPR 2026

We present a scalable 3D reconstruction model that addresses a critical limitation in offline feed-forward methods: their computational and memory requirements grow quadratically w.r.t. the number of input images. Our approach is built on the key insight that this bottleneck stems from the varying-l

Cited by 0SourcecodeScholar
2025

Towards Learning to Complete Anything in Lidar

ICML 2025poster

We propose CAL (Complete Anything in Lidar) for Lidar-based shape-completion in-the-wild. This is closely related to Lidar-based semantic/panoptic scene completion. However, contemporary methods can only complete and recognize objects from a closed vocabulary labeled in existing Lidar datasets. Diff…

Cited by 0SourcePDFScholar
2024

SeMoLi: What Moves Together Belongs Together

CVPR 2024poster

We tackle semi-supervised object detection based on motion cues. Recent results suggest that heuristic-based clustering methods in conjunction with object trackers can be used to pseudo-label instances of moving objects and use these as supervisory signals to train 3D object detectors in Lidar data…

Cited by 8SourcePDFScholar
2023

Walking Your LiDOG: A Journey Through Multiple Domains for LiDAR Semantic Segmentation

ICCV 2023poster

The ability to deploy robots that can operate safely in diverse environments is crucial for developing embodied intelligent agents. As a community, we have made tremendous progress in within-domain LiDAR semantic segmentation. However, do these methods generalize across domains? To answer this que…

Cited by 15PDFcodeScholar
2022

Learning to Discover and Detect Objects

NeurIPS 2022accept

We tackle the problem of novel class discovery and localization (NCDL). In this setting, we assume a source dataset with supervision for only some object classes. Instances of other classes need to be discovered, classified, and localized automatically based on visual similarity without any human su…

2022

Quo Vadis: Is Trajectory Forecasting the Key Towards Long-Term Multi-Object Tracking?

NeurIPS 2022accept

Recent developments in monocular multi-object tracking have been very successful in tracking visible objects and bridging short occlusion gaps, mainly relying on data-driven appearance models. While significant advancements have been made in short-term tracking performance, bridging longer occlusio…

2022

Unsupervised Class-Agnostic Instance Segmentation of 3D LiDAR Data for Autonomous Vehicles

RA-L 2022

Fine-grained scene understanding is essential for autonomous driving. The context around a vehicle can change drastically while navigating, making it hard to identify and understand the different objects that may appear. Although recent efforts on semantic and panoptic segmentation pushed the field

Cited by 26SourceScholar
2021

4D Panoptic LiDAR Segmentation

CVPR 2021poster

Temporal semantic scene understanding is critical for self-driving cars or robots operating in dynamic environments. In this paper, we propose 4D panoptic LiDAR segmentation to assign a semantic class and a temporally-consistent instance ID to a sequence of 3D points. To this end, we present an appr…

Cited by 91PDFcodeScholar
2021

STEP: Segmenting and Tracking Every Pixel

NeurIPS 2021poster

The task of assigning semantic classes and track identities to every pixel in a video is called video panoptic segmentation. Our work is the first that targets this task in a real-world setting requiring dense interpretation in both spatial and temporal domains. As the ground-truth for this task is…

Cited by 89SourcecodeScholar
2020

How to Train Your Deep Multi-Object Tracker

CVPR 2020poster

The recent trend in vision-based multi-object tracking (MOT) is heading towards leveraging the representational power of deep learning to jointly learn to detect and track objects. However, existing methods train only certain sub-modules using loss functions that often do not correlate with establis…

Cited by 274PDFcodeScholar
2020

STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos

ECCV 2020poster

Existing methods for instance segmentation in videos typically involve multi-stage pipelines that follow the tracking-by detection paradigm and model a video clip as a sequence of images. Multiple networks are used to detect objects in individual frames, and then associate these detections over time…

2019

MOTS: Multi-Object Tracking and Segmentation

CVPR 2019poster

This paper extends the popular task of multi-object tracking to multi-object tracking and segmentation (MOTS). Towards this goal, we create dense pixel-level annotations for two existing tracking datasets using a semi-automatic annotation procedure. Our new annotations comprise 65,213 pixel masks fo…

Cited by 699PDFScholar
2016

Multi-scale object candidates for generic object tracking in street scenes

ICRA 2016

Most vision based systems for object tracking in urban environments focus on a limited number of important object categories such as cars or pedestrians, for which powerful detectors are available. However, practical driving scenarios contain many additional objects of interest, for which suitable d

Cited by 44SourceScholar
2015

A fixed-dimensional 3D shape representation for matching partially observed objects in street scenes

ICRA 2015poster

In this paper, we present an object-centric, fixed-dimensional 3D shape representation for robust matching of partially observed object shapes, which is an important component for object categorization from 3D data. A main problem when working with RGB-D data from stereo, Kinect, or laser sensors is…

Cited by 6SourceScholar