← Search

Lorenzo Garattoni

11 accepted papers

2026

Swarm-ReID: Decentralized Self-Adaptive Gallery Construction for Multi-Robot Open-World Person Re-Identification

ICRA 2026poster

Swarm perception enables a robot swarm to collectively sense and understand the environment by integrating sensory inputs from individual robots. We explore its application to person re-identification (re-id), the task of recognizing previously observed individuals. Traditional re-id systems rely on…

Cited by 0Scholar
2025

Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection

CVPR 2025highlight

Weakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-world scenarios. This is due to the fact that RGB-features are not sufficiently distinctive in setting apart categories such…

2025

Mixture of Experts Guided by Gaussian Splatters Matters: A new Approach to Weakly-Supervised Video Anomaly Detection

ICCV 2025poster

Video Anomaly Detection (VAD) is a challenging task due to the variability of anomalous events and the limited availability of labeled data. Under the Weakly-Supervised VAD (WSVAD) paradigm, only video-level labels are provided during training, while predictions are made at the frame level. Although…

2024

HouseCat6D - A Large-Scale Multi-Modal Category Level 6D Object Perception Dataset with Household Objects in Realistic Scenarios

CVPR 2024highlight

Estimating 6D object poses is a major challenge in 3D computer vision. Building on successful instance-level approaches research is shifting towards category-level pose estimation for practical applications. Current category-level datasets however fall short in annotation quality and pose variety. A…

2024

Improving Self-Supervised Learning of Transparent Category Poses With Language Guidance and Implicit Physical Constraints

RA-L 2024

Accurate object pose estimation is crucial for robotic applications and recent trends in category-level pose estimation show great potential for applications encountering a large variety of similar objects, often encountered in home environments. While common in such environments, photometrically ch

Cited by 1SourceScholar
2024

Transferability in the Automatic Off-Line Design of Robot Swarms: From Sim-to-Real to Embodiment and Design-Method Transfer Across Different Platforms

RA-L 2024

Automatic off-line design is an attractive approach to implementing robot swarms. In this approach, a designer specifies a mission to be accomplished by the swarm, and an optimization process generates suitable control software for the individual robots through computer-based simulations. Most relev

Cited by 7SourceScholar
2023

LAC - Latent Action Composition for Skeleton-based Action Segmentation

ICCV 2023poster

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a temporal model to classify frame-wise actions. However, their perfo…

Cited by 14PDFScholar
2023

Self-Supervised Video Representation Learning via Latent Time Navigation

AAAI 2023technical

Self-supervised video representation learning aimed at maximizing similarity between different temporal segments of one video, in order to enforce feature persistence over time. This leads to loss of pertinent information related to temporal relationships, rendering actions such as `enter' and `leav…

Cited by 11SourcePDFScholar
2022

PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation With Photometrically Challenging Objects

CVPR 2022poster

Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape has become a promising trend. As such, a new research field needs to be supported by well-designed datasets. To provide…

Cited by 54PDFScholar
2021

DemoGrasp: Few-Shot Learning for Robotic Grasping with Human Demonstration

IROS 2021poster

The ability to successfully grasp objects is crucial in robotics, as it enables several interactive downstream applications. To this end, most approaches either compute the full 6D pose for the object of interest or learn to predict a set of grasping points. While the former approaches do not scale…

Cited by 38SourceScholar
2019

Toyota Smarthome: Real-World Activities of Daily Living

ICCV 2019poster

The performance of deep neural networks is strongly influenced by the quantity and quality of annotated data. Most of the large activity recognition datasets consist of data sourced from the web, which does not reflect challenges that exist in activities of daily living. In this paper, we introduce…

Cited by 204PDFScholar