← Search

Ales Leonardis

23 accepted papers

2026

Articulation in Motion: Prior-free Part Mobility Analysis for Articulated Objects By Dynamic-Static Disentanglement

ICLR 2026poster

Articulated objects are ubiquitous in daily life. Our goal is to achieve a high-quality reconstruction, segmentation of independent moving parts, and analysis of articulation. Recent methods analyse two different articulation states and perform per-point part segmentation, optimising per-part articu…

Cited by 0SourcecodeScholar
2025

Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics

AAAI 2025technical

With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in re…

Cited by 0SourcePDFScholar
2024

GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose Refinement

CVPR 2024poster

Object pose refinement is essential for robust object pose estimation. Previous work has made significant progress towards instance-level object pose refinement. Yet category-level pose refinement is a more challenging problem due to large shape variations within a category and the discrepancies bet…

Cited by 4SourcePDFScholar
2024

Multi-task Learning with 3D-Aware Regularization

ICLR 2024poster

Deep neural networks have become the standard solution for designing models that can perform multiple dense computer vision tasks such as depth estimation and semantic segmentation thanks to their ability to capture complex correlations in high dimensional feature space across tasks. However, the cr…

2024

bit2bit: 1-bit quanta video reconstruction via self-supervised photon prediction

NeurIPS 2024poster

Quanta image sensors, such as SPAD arrays, are an emerging sensor technology, producing 1-bit arrays representing photon detection events over exposures as short as a few nanoseconds. In practice, raw data are post-processed using heavy spatiotemporal binning to create more useful and interpretable…

2023

Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes

ICCV 2023poster

The success of deep learning models on structured data has generated significant interest in extending their application to non-Euclidean domains. In this work, we introduce a novel intrinsic operator suitable for representation learning on 3D meshes. Our operator is specifically tailored to adapt i…

Cited by 0PDFcodeScholar
2022

End-to-End Learning to Grasp via Sampling From Object Point Clouds

RA-L 2022

The ability to grasp objects is an essential skill that enables many robotic manipulation tasks. Recent works have studied point cloud-based methods for object grasping by starting from simulated datasets and have shown promising performance in real-world scenarios. Nevertheless, many of them still

Cited by 35SourcecodeScholar
2022

Model-Based Image Signal Processors via Learnable Dictionaries

AAAI 2022technical

Digital cameras transform sensor RAW readings into RGB images by means of their Image Signal Processor (ISP). Computational photography tasks such as image denoising and colour constancy are commonly performed in the RAW domain, in part due to the inherent hardware design, but also due to the appeal…

2022

Residual Contrastive Learning for Image Reconstruction: Learning Transferable Representations from Noisy Images

IJCAI 2022poster

This paper is concerned with contrastive learning (CL) for low-level image restoration and enhancement tasks. We propose a new label-efficient learning paradigm based on residuals, residual contrastive learning (RCL), and derive an unsupervised visual representation learning framework, suitable for…

Cited by 5SourcePDFScholar
2021

FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation Mechanism

CVPR 2021poster

In this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net)…

Cited by 197PDFcodeScholar
2020

A Multi-Hypothesis Approach to Color Constancy

CVPR 2020poster

Contemporary approaches frame the color constancy problem as learning camera specific illuminant mappings. While high accuracy can be achieved on camera specific data, these models depend on camera spectral sensitivity and typically exhibit poor generalisation to new devices. Additionally, regressio…

Cited by 65PDFScholar
2020

G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector Features

CVPR 2020poster

In this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divide-and-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D…

Cited by 133PDFcodeScholar
2020

More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning

ECCV 2020poster

Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning","Overcoming catastrophic forgetting in neural networks is a long-standing and core research objective for incremental learning. Notable studies have shown regularization strategies enable the network to remember previousl…

2020

Unsupervised Model Personalization While Preserving Privacy and Scalability: An Open Problem

CVPR 2020poster

This work investigates the task of unsupervised model personalization, adapted to continually evolving, unlabeled local user images. We consider the practical scenario where a high capacity server interacts with a myriad of resource-limited edge devices, imposing strong requirements on scalability a…

Cited by 35PDFcodeScholar
2018

Learning to Exploit Stability for 3D Scene Parsing

NeurIPS 2018poster

Human scene understanding uses a variety of visual and non-visual cues to perform inference on object types, poses, and relations. Physics is a rich and universal cue which we exploit to enhance scene understanding. We integrate the physical cue of stability into the learning process using a REINFOR…

Cited by 49SourcePDFScholar
2017

Beyond Standard Benchmarks: Parameterizing Performance Evaluation in Visual Object Tracking

ICCV 2017poster

Object-to-camera motion produces a variety of apparent motion patterns that significantly affect performance of short-term visual trackers. Despite being crucial for designing robust trackers, their influence is poorly explored in standard benchmarks due to weakly defined, biased and overlapping att…

Cited by 26PDFScholar
2016

Hierarchical spatial model for 2D range data based room categorization

ICRA 2016

The next generation service robots are expected to co-exist with humans in their homes. Such a mobile robot requires an efficient representation of space, which should be compact and expressive, for effective operation in real-world environments. In this paper we present a novel approach for 2D grou

Cited by 5SourceScholar
2016

Task-relevant grasp selection: A joint solution to planning grasps and manipulative motion trajectories

IROS 2016poster

This paper addresses the problem of jointly planning both grasps and subsequent manipulative actions. Previously, these two problems have typically been studied in isolation, however joint reasoning is essential to enable robots to complete real manipulative tasks. In this paper, the two problems ar…

Cited by 25SourceScholar
2015

Compositional Hierarchical Representation of Shape Manifolds for Classification of Non-Manifold Shapes

ICCV 2015poster

We address the problem of statistical learning of shape models which are invariant to translation, rotation and scale in compositional hierarchies when data spaces of measurements and shape spaces are not topological manifolds. In practice, this problem is observed while modeling shapes having multi…

Cited by 7PDFScholar
2015

Single Target Tracking Using Adaptive Clustered Decision Trees and Dynamic Multi-Level Appearance Models

CVPR 2015poster

This paper presents a method for single target tracking of arbitrary objects in challenging video sequences. Targets are modeled at three different levels of granularity (pixel level, parts-based level and bounding box level), which are cross-constrained to enable robust model relearning. The main c…

Cited by 71SourcePDFScholar