← Search

Davide Moltisanti

8 accepted papers

2026

ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos

CVPR 2026

Procedural planning aims to predict a sequence of actions that transforms an initial visual state into a desired goal, a fundamental ability for intelligent agents operating in complex environments. Existing approaches typically rely on large-scale models that learn procedural structures implicitly,

Cited by 0SourceScholar
2025

HD-EPIC: A Highly-Detailed Egocentric Video Dataset

CVPR 2025poster

We present a validation dataset of newly-collected kitchen based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, al…

Cited by 3SourcePDFScholar
2024

Efficient Pre-training for Localized Instruction Generation of Procedural Videos

ECCV 2024poster

"Procedural videos, exemplified by recipe demonstrations, are instrumental in conveying step-by-step instructions. However, understanding such videos is challenging as it involves the precise localization of steps and the generation of textual instructions. Manually annotating steps and writing inst…

2023

Learning Action Changes by Measuring Verb-Adverb Textual Relationships

CVPR 2023poster

The goal of this work is to understand the way actions are performed in videos. That is, given a video, we aim to predict an adverb indicating a modification applied to the action (e.g. cut "finely"). We cast this problem as a regression task. We measure textual relationships between verbs and adver…

2022

BRACE: The Breakdancing Competition Dataset for Dance Motion Synthesis

ECCV 2022poster

"Generative models for audio-conditioned dance motion synthesis map music features to dance movements. Models are trained to associate motion patterns to audio patterns, usually without an explicit knowledge of the human body. This approach relies on a few assumptions: strong music-dance correlation…

2019

Action Recognition From Single Timestamp Supervision in Untrimmed Videos

CVPR 2019poster

Recognising actions in videos relies on labelled supervision during training, typically the start and end times of each action instance. This supervision is not only subjective, but also expensive to acquire. Weak video-level supervision has been successfully exploited for recognition in untrimmed v…

Cited by 92PDFcodeScholar
2018

Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

ECCV 2018poster

First-person vision is gaining interest as it offers a unique viewpoint on people’s interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due to the lack of sufficiently large datasets. In this paper, we introduce EPIC-KI…

Cited by 1329SourcePDFScholar
2017

Trespassing the Boundaries: Labeling Temporal Bounds for Object Interactions in Egocentric Video

ICCV 2017poster

Manual annotations of temporal bounds for object interactions (i.e. start and end times) are typical training input to recognition, localization and detection algorithms. For three publicly available egocentric datasets, we uncover inconsistencies in ground truth temporal bounds within and across an…

Cited by 36PDFScholar