← Search

Slobodan Ilic

35 accepted papers

2026

CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis

ICLR 2026poster

Spectral imaging offers promising applications across diverse domains, including medicine and urban scene understanding, and is already established as a critical modality in remote sensing. However, variability in channel dimensionality and captured wavelengths among spectral cameras impede the deve…

Cited by 0SourcecodeScholar
2025

Multi-Layer Feature Exchange Transformer for Multi-View 6D Object Pose Estimation in Robot Bin Picking

ICRA 2025

Accurate 6D object pose estimation is crucial in industrial automation, particularly in robotic bin picking, where objects are often textureless, reflective, and arranged in cluttered environments. Multi-view pose estimation methods offer significant advantages over single-view methods by providing

Cited by 0SourceScholar
2025

RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation

ICCV 2025poster

Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose predictions. To address this, we reformulate template-based object pose…

Cited by 0SourcePDFScholar
2025

SeaLion: Semantic Part-Aware Latent Point Diffusion Models for 3D Generation

CVPR 2025poster

Denoising diffusion probabilistic models have achieved significant success in point cloud generation, enabling numerous downstream applications, such as generative data augmentation and 3D model editing. However, little attention has been given to generating point clouds with point-wise segmentation…

Cited by 0SourcePDFScholar
2025

Spiral: Semantic-Aware Progressive LiDAR Scene Generation and Understanding

NeurIPS 2025poster

Leveraging diffusion models, 3D LiDAR scene generation has achieved great success in both range-view and voxel-based representations. While recent voxel-based approaches can generate both geometric structures and semantic labels, existing range-view methods are limited to producing unlabeled LiDAR s…

Cited by 0SourceScholar
2024

MatchU: Matching Unseen Objects for 6D Pose Estimation from RGB-D Images

CVPR 2024poster

Recent learning methods for object pose estimation require resource-intensive training for each individual object instance or category hampering their scalability in real applications when confronted with previously unseen objects. In this paper we propose MatchU a Fuse-Describe-Match strategy for 6…

Cited by 10SourcePDFScholar
2023

On the Importance of Accurate Geometry Data for Dense 3D Vision Tasks

CVPR 2023poster

Learning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks. These are typically not compared nor discussed in the literature due to a lack of multi-modal datasets. Texture-less r…

2023

Rotation-Invariant Transformer for Point Cloud Matching

CVPR 2023poster

The intrinsic rotation invariance lies at the core of matching point clouds with handcrafted descriptors. However, it is widely despised by recent deep matchers that obtain the rotation invariance extrinsically via data augmentation. As the finite number of augmented rotations can never span the con…

2022

PolarMesh: A Star-Convex 3D Shape Approximation for Object Pose Estimation

RA-L 2022

In this letter, we introduce PolarMesh as a star-convex approximation of a 3D object based on spherical projection and can be applied to monocular object pose and shape estimation. The proposed PolarMesh can be stored in a discrete 2D map that allows a trivial conversion between it and the object su

Cited by 11SourceScholar
2022

WS-OPE: Weakly Supervised 6-D Object Pose Regression Using Relative Multi-Camera Pose Constraints

RA-L 2022

Precise annotation of 6-D poses in real data is intricate and time-consuming, however, an essential requirement to train pose estimation pipelines. We propose a way for scalable, end-to-end 6-D pose regression with weak supervision to avoid this problem. Our method requires neither 3-D models nor 6-

Cited by 11SourceScholar
2022

WeLSA: Learning to Predict 6D Pose from Weakly Labeled Data Using Shape Alignment

ECCV 2022poster

"Object pose estimation is a crucial task in computer vision and augmented reality. One of its key challenges is the difficulty of annotation of real training data and the lack of textured CAD models. Therefore, pipelines which do not require CAD models and which can be trained with few labeled imag…

Cited by 5SourcePDFScholar
2021

CoFiNet: Reliable Coarse-to-fine Correspondences for Robust PointCloud Registration

NeurIPS 2021poster

We study the problem of extracting correspondences between a pair of point clouds for registration. For correspondence retrieval, existing works benefit from matching sparse keypoints detected from dense points but usually struggle to guarantee their repeatability. To address this issue, we present…

2020

6D Camera Relocalization in Ambiguous Scenes via Continuous Multimodal Inference

ECCV 2020poster

We present a multimodal camera relocalization framework that captures ambiguities and uncertainties with continuous mixture models defined on the manifold of camera poses. In highly ambiguous environments, which can easily arise due to symmetries and repetitive structures in the scene, computing one…

2019

Seeing Beyond Appearance - Mapping Real Images into Geometrical Domains for Unsupervised CAD-based Recognition

IROS 2019poster

While convolutional neural networks are dominating the field of computer vision, one usually does not have access to the large amount of domain-relevant data needed for their training. Therefore, it has become common practice to use available synthetic samples along domain adaptation schemes to prep…

Cited by 14SourceScholar
2018

A Minimalist Approach to Type-Agnostic Detection of Quadrics in Point Clouds

CVPR 2018poster

This paper proposes a segmentation-free, automatic and efficient procedure to detect general geometric quadric forms in point clouds, where clutter and occlusions are inevitable. Our everyday world is dominated by man-made objects which are designed using 3D primitives (such as planes, cones, sphere…

Cited by 15SourcePDFScholar
2018

Bayesian Pose Graph Optimization via Bingham Distributions and Tempered Geodesic MCMC

NeurIPS 2018poster

We introduce Tempered Geodesic Markov Chain Monte Carlo (TG-MCMC) algorithm for initializing pose graph optimization problems, arising in various scenarios such as SFM (structure from motion) or SLAM (simultaneous localization and mapping). TG-MCMC is first of its kind as it unites global non-convex…

Cited by 42SourcePDFScholar
2018

PPF-FoldNet: Unsupervised Learning of Rotation Invariant 3D Local Descriptors

ECCV 2018poster

We present PPF-FoldNet for unsupervised learning of 3D local descriptors on pure point cloud geometry. Based on the folding-based auto-encoding of well known point pair features, PPF-FoldNet offers many desirable properties: it necessitates neither supervision, nor a sensitive local reference frame,…

Cited by 476SourcePDFScholar
2018

PPFNet: Global Context Aware Local Features for Robust 3D Point Matching

CVPR 2018poster

We present PPFNet - Point Pair Feature NETwork for deeply learning a globally informed 3D local feature descriptor to find correspondences in unorganized point clouds. PPFNet learns local descriptors on pure geometry and is highly aware of the global context, an important cue in deep learning. Our 3…

Cited by 764SourcePDFScholar
2018

SobolevFusion: 3D Reconstruction of Scenes Undergoing Free Non-Rigid Motion

CVPR 2018poster

We present a system that builds 3D models of non-rigidly moving surfaces from scratch in real time using a single RGB-D stream. Our solution is based on the variational level set method, thus it copes with arbitrary geometry, including topological changes. It warps a given truncated signed distance…

Cited by 119SourcePDFScholar
2018

When Regression Meets Manifold Learning for Object Recognition and Pose Estimation

ICRA 2018poster

In this work, we propose a method for object recognition and pose estimation from depth images using convolutional neural networks. Previous methods addressing this problem rely on manifold learning to learn low dimensional viewpoint descriptors and employ them in a nearest neighbor search on an est…

Cited by 35SourceScholar
2017

3D object instance recognition and pose estimation using triplet loss with dynamic margin

IROS 2017poster

In this paper, we address the problem of 3D object instance recognition and pose estimation of localized objects in cluttered environments using convolutional neural networks. Inspired by the descriptor learning approach of Wohlhart et al. [1], we propose a method that introduces the dynamic margin…

Cited by 48SourceScholar
2017

KillingFusion: Non-Rigid 3D Reconstruction Without Correspondences

CVPR 2017spotlight

We introduce a geometry-driven approach for real-time 3D reconstruction of deforming surfaces from a single RGB-D stream without any templates or shape priors. To this end, we tackle the problem of non-rigid registration by level set evolution without explicit correspondence search. Given a pair of…

Cited by 210PDFScholar
2017

Real-Time 3D Model Tracking in Color and Depth on a Single CPU Core

CVPR 2017poster

We present a novel method to track 3D models in color and depth data. To this end, we introduce approximations that accelerate the state-of-the-art in region-based tracking by an order of magnitude while retaining similar accuracy. Furthermore, we show how the method can be made more robust in the p…

Cited by 51PDFScholar
2017

SSD-6D: Making RGB-Based 3D Detection and 6D Pose Estimation Great Again

ICCV 2017oral

We present a novel method for detecting 3D model instances and estimating their 6D poses from RGB data in a single shot. To this end, we extend the popular SSD paradigm to cover the full 6D pose space and train on synthetic model data only. Our approach competes or surpasses current state-of-the-art…

Cited by 1284PDFScholar
2016

Volumetric 3D Tracking by Detection

CVPR 2016spotlight

In this paper, we propose a new framework for 3D tracking by detection based on fully volumetric representations. On one hand, 3D tracking by detection has shown robust use in the context of interaction (Kinect) and surface tracking. On the other hand, volumetric representations have recently been p…

Cited by 40PDFScholar
2015

A Versatile Learning-Based 3D Temporal Tracker: Scalable, Robust, Online

ICCV 2015poster

This paper proposes a temporal tracking algorithm based on Random Forest that uses depth images to estimate and track the 3D pose of a rigid object in real-time. Compared to the state of the art aimed at the same goal, our algorithm holds important attributes such as high robustness against holes an…

Cited by 87PDFScholar
2015

Toward User-Specific Tracking by Detection of Human Shapes in Multi-Cameras

CVPR 2015poster

Human shape tracking consists in fitting a template model to temporal sequences of visual observations. It usually comprises an association step, that finds correspondences between the model and the input data, and a deformation step, that fits the model to the observations given correspondences. Mo…

Cited by 19SourcePDFScholar