← Search

Gianpiero Francesca

12 accepted papers

2026

MoVie: Broaden Your Views with Human Motion for Action Detection

CVPR 2026

Human action detection in videos requires both semantic recognition and accurate modeling of motion. While recent video foundation models have advanced visual semantics, they still struggle to capture complex and compositional actions due to the limited representation ability of motion. Human skelet

Cited by 0SourceScholar
2026

Swarm-ReID: Decentralized Self-Adaptive Gallery Construction for Multi-Robot Open-World Person Re-Identification

ICRA 2026poster

Swarm perception enables a robot swarm to collectively sense and understand the environment by integrating sensory inputs from individual robots. We explore its application to person re-identification (re-id), the task of recognizing previously observed individuals. Traditional re-id systems rely on…

Cited by 0Scholar
2025

Hierarchical Vector Quantization for Unsupervised Action Segmentation

AAAI 2025technical

In this work, we address unsupervised temporal action segmentation, which segments a set of long, untrimmed videos into semantically meaningful segments that are consistent across videos. While recent approaches combine representation learning and clustering in a single step for this task, they do n…

2025

Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection

CVPR 2025highlight

Weakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-world scenarios. This is due to the fact that RGB-features are not sufficiently distinctive in setting apart categories such…

2025

MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation

CVPR 2025poster

Long-term dense action anticipation is very challenging since it requires predicting actions and their durations several minutes into the future based on provided video observations. To model the uncertainty of future outcomes, stochastic models predict several potential future action sequences for…

2025

Mixture of Experts Guided by Gaussian Splatters Matters: A new Approach to Weakly-Supervised Video Anomaly Detection

ICCV 2025poster

Video Anomaly Detection (VAD) is a challenging task due to the variability of anomalous events and the limited availability of labeled data. Under the Weakly-Supervised VAD (WSVAD) paradigm, only video-level labels are provided during training, while predictions are made at the frame level. Although…

2024

Gated Temporal Diffusion for Stochastic Long-term Dense Anticipation

ECCV 2024poster

"Long-term action anticipation has become an important task for many applications such as autonomous driving and human-robot interaction. Unlike short-term anticipation, predicting more actions into the future imposes a real challenge with the increasing uncertainty in longer horizons. While there h…

2024

Transferability in the Automatic Off-Line Design of Robot Swarms: From Sim-to-Real to Embodiment and Design-Method Transfer Across Different Platforms

RA-L 2024

Automatic off-line design is an attractive approach to implementing robot swarms. In this approach, a designer specifies a mission to be accomplished by the swarm, and an optimization process generates suitable control software for the individual robots through computer-based simulations. Most relev

Cited by 7SourceScholar
2023

How Much Temporal Long-Term Context is Needed for Action Segmentation?

ICCV 2023poster

Modeling long-term context in videos is crucial for many fine-grained tasks including temporal action segmentation. An interesting question that is still open is how much long-term temporal context is needed for optimal performance. While transformers can model the long-term context of a video, this…

Cited by 39PDFcodeScholar
2023

LAC - Latent Action Composition for Skeleton-based Action Segmentation

ICCV 2023poster

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a temporal model to classify frame-wise actions. However, their perfo…

Cited by 14PDFScholar
2023

Self-Supervised Video Representation Learning via Latent Time Navigation

AAAI 2023technical

Self-supervised video representation learning aimed at maximizing similarity between different temporal segments of one video, in order to enforce feature persistence over time. This leads to loss of pertinent information related to temporal relationships, rendering actions such as `enter' and `leav…

Cited by 11SourcePDFScholar
2019

Toyota Smarthome: Real-World Activities of Daily Living

ICCV 2019poster

The performance of deep neural networks is strongly influenced by the quantity and quality of annotated data. Most of the large activity recognition datasets consist of data sourced from the web, which does not reflect challenges that exist in activities of daily living. In this paper, we introduce…

Cited by 204PDFScholar