← Search

Laura Sevilla-Lara

16 accepted papers

2026

Mask2IV: Interaction-Centric Video Generation via Mask Trajectories

AAAI 2026technical

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training, and affordance reasoning. However, existing methods often s

Cited by 0SourcePDFScholar
2025

Learning Precise Affordances from Egocentric Videos for Robotic Manipulation

ICCV 2025poster

Affordance, defined as the potential actions that an object offers, is crucial for embodied AI agents. For example, such knowledge directs an agent to grasp a knife by the handle for cutting or by the blade for safe handover. While existing approaches have made notable progress, affordance research…

2025

Predicting Implicit Arguments in Procedural Video Instructions

ACL 2025long

Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by identifying predicate-argument structure like verb,what,where/with. Procedural instructions are highly elliptic, for insta…

2025

Principles of Visual Tokens for Efficient Video Understanding

ICCV 2025poster

Video understanding has made huge strides in recent years, relying largely on the power of transformers. As this architecture is notoriously expensive and video data is highly redundant, research into improving efficiency has become particularly relevant. Some creative solutions include token select…

2025

Progressive Data Dropout: An Embarrassingly Simple Approach to Train Faster

NeurIPS 2025poster

The success of the machine learning field has reliably depended on training on large datasets. While effective, this trend comes at an extraordinary cost. This is due to two deeply intertwined factors: the size of models and the size of datasets. While promising research efforts focus on reducing th…

Cited by 0SourcecodeScholar
2024

Efficient Pre-training for Localized Instruction Generation of Procedural Videos

ECCV 2024poster

"Procedural videos, exemplified by recipe demonstrations, are instrumental in conveying step-by-step instructions. However, understanding such videos is challenging as it involves the precise localization of steps and the generation of textual instructions. Manually annotating steps and writing inst…

2024

One-Shot Open Affordance Learning with Foundation Models

CVPR 2024poster

We introduce One-shot Open Affordance Learning (OOAL) where a model is trained with just one example per base object category but is expected to identify novel objects and affordances. While vision-language models excel at recognizing novel objects and scenes they often struggle to understand finer…

2023

LOCATE: Localize and Transfer Object Parts for Weakly Supervised Affordance Grounding

CVPR 2023poster

Humans excel at acquiring knowledge through observation. For example, we can learn to use new tools by watching demonstrations. This skill is fundamental for intelligent systems to interact with the world. A key step to acquire this skill is to identify what part of the object affords each action, w…

Cited by 51SourcePDFScholar
2023

Learning Action Changes by Measuring Verb-Adverb Textual Relationships

CVPR 2023poster

The goal of this work is to understand the way actions are performed in videos. That is, given a video, we aim to predict an adverb indicating a modification applied to the action (e.g. cut "finely"). We cast this problem as a regression task. We measure textual relationships between verbs and adver…

2022

CLASTER: Clustering with Reinforcement Learning for Zero-Shot Action Recognition

ECCV 2022poster

"Zero-Shot action recognition is the task of recognizing action classes without visual examples. The problem can be seen as learning a representation on seen classes which generalizes well to instances of unseen classes, without losing discriminability between classes. Neural networks are able to mo…

Cited by 41SourcePDFScholar
2022

Learn2Augment: Learning to Composite Videos for Data Augmentation in Action Recognition

ECCV 2022poster

"We address the problem of data augmentation for video action recognition. Standard augmentation strategies in video are hand designed and sample the space of possible augmented data points either at random, without knowing which augmented points will be better, or through heuristics. We propose to…

Cited by 49SourcePDFScholar
2021

Adaptive Prototype Learning and Allocation for Few-Shot Segmentation

CVPR 2021poster

Prototype learning is extensively used for few-shot segmentation. Typically, a single prototype is obtained from the support feature by averaging the global object information. However, using one prototype to represent all the information may lead to ambiguities. In this paper, we propose two novel…

Cited by 473PDFcodeScholar
2019

DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition

CVPR 2019poster

Motion has shown to be useful for video understanding, where motion is typically represented by optical flow. However, computing flow from video frames is very timeconsuming. Recent works directly leverage the motion vectors and residuals readily available in the compressed video to represent motion…

Cited by 168PDFScholar
2016

Optical Flow With Semantic Segmentation and Localized Layers

CVPR 2016spotlight

Existing optical flow methods make generic, spatially homogeneous, assumptions about the spatial structure of the flow. In reality, optical flow varies across an image depending on object class. Simply put, different objects move differently. Here we exploit recent advances in static semantic scene…

Cited by 251PDFScholar