← Search

Laura Leal-Taixé

50 accepted papers

2026

Motion Attribution for Video Generation

ICML 2026oral

Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework that scales to modern, large, high-quality video datasets and m…

Cited by 2SourceScholar
2026

Scene-Centric Unsupervised Video Panoptic Segmentation

CVPR 2026

Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision. Existing unsupervised scene understanding works mainly focuse

Cited by 0SourceScholar
2026

VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale

CVPR 2026

We present a scalable 3D reconstruction model that addresses a critical limitation in offline feed-forward methods: their computational and memory requirements grow quadratically w.r.t. the number of input images. Our approach is built on the key insight that this bottleneck stems from the varying-l

Cited by 0SourcecodeScholar
2025

Towards Learning to Complete Anything in Lidar

ICML 2025poster

We propose CAL (Complete Anything in Lidar) for Lidar-based shape-completion in-the-wild. This is closely related to Lidar-based semantic/panoptic scene completion. However, contemporary methods can only complete and recognize objects from a closed vocabulary labeled in existing Lidar datasets. Diff…

Cited by 0SourcePDFScholar
2025

Towards a Unified Copernicus Foundation Model for Earth Vision

ICCV 2025poster

Advances in Earth observation (EO) foundation models have unlocked the potential of big satellite data to learn generic representations from space, benefiting a wide range of downstream applications crucial to our planet. However, most existing efforts remain limited to fixed spectral sensors, focus…

2024

Better Call SAL: Towards Learning to Segment Anything in Lidar

ECCV 2024poster

"We propose the (Segment Anything in Lidar) method consisting of a text-promptable zero-shot model for segmenting and classifying any object in Lidar, and a pseudo-labeling engine that facilitates model training without manual supervision. While the established paradigm for (LPS) relies on manual su…

2024

MICDrop: Masking Image and Depth Features via Complementary Dropout for Domain-Adaptive Semantic Segmentation

ECCV 2024poster

"Unsupervised Domain Adaptation (UDA) is the task of bridging the domain gap between a labeled source domain, e.g., synthetic data, and an unlabeled target domain. We observe that current UDA methods show inferior results on fine structures and tend to oversegment objects with ambiguous appearance.…

2024

SPAMming Labels: Efficient Annotations for the Trackers of Tomorrow

ECCV 2024poster

"Increasing the annotation efficiency of trajectory annotations from videos has the potential to enable the next generation of data-hungry tracking algorithms to thrive on large-scale datasets. Despite the importance of this task, there are currently very few works exploring how to efficiently label…

Cited by 1SourcePDFScholar
2024

SatSynth: Augmenting Image-Mask Pairs through Diffusion Models for Aerial Semantic Segmentation

CVPR 2024poster

In recent years semantic segmentation has become a pivotal tool in processing and interpreting satellite imagery. Yet a prevalent limitation of supervised learning techniques remains the need for extensive manual annotations by experts. In this work we explore the potential of generative image diffu…

Cited by 23SourcePDFScholar
2024

The Nerfect Match: Exploring NeRF Features for Visual Localization

ECCV 2024poster

"In this work, we propose the use of Neural Radiance Fields () as a scene representation for visual localization. Recently, has been employed to enhance pose regression and scene coordinate regression models by augmenting the training database, providing auxiliary supervision through rendered images…

Cited by 9SourcePDFScholar
2023

G-MSM: Unsupervised Multi-Shape Matching With Graph-Based Affinity Priors

CVPR 2023poster

We present G-MSM (Graph-based Multi-Shape Matching), a novel unsupervised learning approach for non-rigid shape correspondence. Rather than treating a collection of input poses as an unordered set of samples, we explicitly model the underlying shape data manifold. To this end, we propose an adaptive…

2023

Lidar Panoptic Segmentation and Tracking without Bells and Whistles

IROS 2023poster

State-of-the-art lidar panoptic segmentation (LPS) methods follow “bottom-up” segmentation-centric fashion wherein they build upon semantic segmentation networks by utilizing clustering to obtain object instances. In this paper, we re-think this approach and propose a surprisingly simple yet effecti…

Cited by 8SourcecodeScholar
2023

Simple Cues Lead to a Strong Multi-Object Tracker

CVPR 2023poster

For a long time, the most common paradigm in MultiObject Tracking was tracking-by-detection (TbD), where objects are first detected and then associated over video frames. For association, most models resourced to motion and appearance cues, e.g., re-identification networks. Recent approaches based o…

2023

Soft Augmentation for Image Classification

CVPR 2023poster

Modern neural networks are over-parameterized and thus rely on strong regularization such as data augmentation and weight decay to reduce overfitting and improve generalization. The dominant form of data augmentation applies invariant transforms, where the learning target of a sample is invariant to…

2023

Unifying Short and Long-Term Tracking With Graph Hierarchies

CVPR 2023poster

Tracking objects over long videos effectively means solving a spectrum of problems, from short-term association for un-occluded objects to long-term association for objects that are occluded and then reappear in the scene. Methods tackling these two tasks are often disjoint and crafted for specific…

2023

Walking Your LiDOG: A Journey Through Multiple Domains for LiDAR Semantic Segmentation

ICCV 2023poster

The ability to deploy robots that can operate safely in diverse environments is crucial for developing embodied intelligent agents. As a community, we have made tremendous progress in within-domain LiDAR semantic segmentation. However, do these methods generalize across domains? To answer this que…

Cited by 15PDFcodeScholar
2022

A Unified Framework for Implicit Sinkhorn Differentiation

CVPR 2022poster

The Sinkhorn operator has recently experienced a surge of popularity in computer vision and related fields. One major reason is its ease of integration into deep learning frameworks. To allow for an efficient training of respective neural networks, we propose an algorithm that obtains analytical gra…

Cited by 26PDFcodeScholar
2022

DirectTracker: 3D Multi-Object Tracking Using Direct Image Alignment and Photometric Bundle Adjustment

IROS 2022poster

Direct methods have shown excellent performance in the applications of visual odometry and SLAM. In this work we propose to leverage their effectiveness for the task of 3D multi-object tracking. To this end, we propose DirectTracker, a framework that effectively combines direct image alignment for t…

Cited by 5SourceScholar
2022

DynamicEarthNet: Daily Multi-Spectral Satellite Dataset for Semantic Change Segmentation

CVPR 2022poster

Earth observation is a fundamental tool for monitoring the evolution of land use in specific areas of interest. Observing and precisely defining change, in this context, requires both time-series data and pixel-wise segmentations. To that end, we propose the DynamicEarthNet dataset that consists of…

Cited by 107PDFScholar
2022

Forecasting From LiDAR via Future Object Detection

CVPR 2022poster

Object detection and forecasting are fundamental components of embodied perception. These two problems, however, are largely studied in isolation by the community. In this paper, we propose an end-to-end approach for motion forecasting based on raw sensor measurement as opposed to ground truth track…

Cited by 40PDFcodeScholar
2022

Is Geometry Enough for Matching in Visual Localization?

ECCV 2022poster

"In this paper, we propose to go beyond the well-established approach to vision-based localization that relies on visual descriptor matching between a query image and a 3D point cloud. While matching keypoints via visual descriptors makes localization highly accurate, it has significant storage dema…

2022

Learning to Discover and Detect Objects

NeurIPS 2022accept

We tackle the problem of novel class discovery and localization (NCDL). In this setting, we assume a source dataset with supervision for only some object classes. Instances of other classes need to be discovered, classified, and localized automatically based on visual similarity without any human su…

2022

Not All Labels Are Equal: Rationalizing the Labeling Costs for Training Object Detection

CVPR 2022poster

Deep neural networks have reached high accuracy on object detection but their success hinges on large amounts of labeled data. To reduce the labels dependency, various active learning strategies have been proposed, typically based on the confidence of the detector. However, these methods are biased…

Cited by 50PDFcodeScholar
2022

PolarMOT: How Far Can Geometric Relations Take Us in 3D Multi-Object Tracking?

ECCV 2022poster

"Most (3D) multi-object tracking methods rely on appearance-based cues for data association. By contrast, we investigate how far we can get by only encoding geometric relationships between objects in 3D space as cues for data-driven data association. We encode 3D detections as nodes in a graph, wher…

Cited by 57SourcePDFScholar
2022

Quo Vadis: Is Trajectory Forecasting the Key Towards Long-Term Multi-Object Tracking?

NeurIPS 2022accept

Recent developments in monocular multi-object tracking have been very successful in tracking visible objects and bridging short occlusion gaps, mainly relying on data-driven appearance models. While significant advancements have been made in short-term tracking performance, bridging longer occlusio…

2022

The Unreasonable Effectiveness of Fully-Connected Layers for Low-Data Regimes

NeurIPS 2022accept

Convolutional neural networks were the standard for solving many computer vision tasks until recently, when Transformers of MLP-based architectures have started to show competitive performance. These architectures typically have a vast number of weights and need to be trained on massive datasets; he…

Cited by 8SourcePDFScholar
2022

TrackFormer: Multi-Object Tracking With Transformers

CVPR 2022poster

The challenging task of multi-object tracking (MOT) requires simultaneous reasoning about track initialization, identity, and spatio-temporal trajectories. We formulate this task as a frame-to-frame set prediction problem and introduce TrackFormer, an end-to-end trainable MOT approach based on an en…

Cited by 1011PDFcodeScholar
2022

Unsupervised Class-Agnostic Instance Segmentation of 3D LiDAR Data for Autonomous Vehicles

RA-L 2022

Fine-grained scene understanding is essential for autonomous driving. The context around a vehicle can change drastically while navigating, making it hard to identify and understand the different objects that may appear. Although recent efforts on semantic and panoptic segmentation pushed the field

Cited by 26SourceScholar
2021

(Just) A Spoonful of Refinements Helps the Registration Error Go Down

ICCV 2021poster

In this paper, we tackle data-driven 3D point cloud registration. Given point correspondences, the standard Kabsch algorithm provides an optimal rotation estimate. This allows to train registration models in an end-to-end manner by differentiating the SVD operation. However, given the initial rotati…

Cited by 3PDFcodeScholar
2021

DENETHOR: The DynamicEarthNET dataset for Harmonized, inter-Operable, analysis-Ready, daily crop monitoring from space

NeurIPS 2021poster

Recent advances in remote sensing products allow near-real time monitoring of the Earth’s surface. Despite increasing availability of near-daily time-series of satellite imagery, there has been little exploration of deep learning methods to utilize the unprecedented temporal density of observations.…

Cited by 54SourceScholar
2021

Learning Intra-Batch Connections for Deep Metric Learning

ICML 2021spotlight

The goal of metric learning is to learn a function that maps samples to a lower-dimensional space where similar samples lie closer than dissimilar ones. Particularly, deep metric learning utilizes neural networks to learn such a mapping. Most approaches rely on losses that only take the relations be…

2021

MG-GAN: A Multi-Generator Model Preventing Out-of-Distribution Samples in Pedestrian Trajectory Prediction

ICCV 2021poster

Pedestrian trajectory prediction is challenging due to its uncertain and multimodal nature. While generative adversarial networks can learn a distribution over future trajectories, they tend to predict out-of-distribution samples when the distribution of future trajectories is a mixture of multiple,…

Cited by 143PDFcodeScholar
2021

MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?

ICCV 2021poster

Deep learning-based methods for video pedestrian detection and tracking require large volumes of training data to achieve good performance. However, data acquisition in crowded public environments raises data privacy concerns -- we are not allowed to simply record and store data without the explicit…

Cited by 159PDFScholar
2021

STEP: Segmenting and Tracking Every Pixel

NeurIPS 2021poster

The task of assigning semantic classes and track identities to every pixel in a video is called video panoptic segmentation. Our work is the first that targets this task in a real-world setting requiring dense interpretation in both spatial and temporal domains. As the ground-truth for this task is…

Cited by 89SourcecodeScholar
2021

The Center of Attention: Center-Keypoint Grouping via Attention for Multi-Person Pose Estimation

ICCV 2021poster

We introduce CenterGroup, an attention-based framework to estimate human poses from a set of identity-agnostic keypoints and person center predictions in an image. Our approach uses a transformer to obtain context-aware embeddings for all detected keypoints and centers and then applies multi-head at…

Cited by 59PDFcodeScholar
2021

Vision-Based Mobile Robotics Obstacle Avoidance With Deep Reinforcement Learning

ICRA 2021poster

Obstacle avoidance is a fundamental and challenging problem for autonomous navigation of mobile robots. In this paper, we consider the problem of obstacle avoidance in simple 3D environments where the robot has to solely rely on a single monocular camera. In particular, we are interested in solving…

Cited by 58SourceScholar
2020

Deep Shells: Unsupervised Shape Correspondence with Optimal Transport

NeurIPS 2020poster

We propose a novel unsupervised learning approach to 3D shape correspondence that builds a multiscale matching pipeline into a deep neural network. This approach is based on smooth shells, the current state-of-the-art axiomatic correspondence method, which requires an a priori stochastic search over…

2020

STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos

ECCV 2020poster

Existing methods for instance segmentation in videos typically involve multi-stage pipelines that follow the tracking-by detection paradigm and model a video clip as a sequence of images. Multiple networks are used to detect objects in individual frames, and then associate these detections over time…

2020

The Group Loss for Deep Metric Learning

ECCV 2020poster

Deep metric learning has yielded impressive results in tasks such as clustering and image retrieval by leveraging neural networks to obtain highly discriminative feature embeddings, which can be used to group samples into different classes. Much research has been devoted to the design of smart loss…

2020

To Learn or Not to Learn: Visual Localization from Essential Matrices

ICRA 2020poster

Visual localization is the problem of estimating a camera within a scene and a key technology for autonomous robots. State-of-the-art approaches for accurate visual localization use scene-specific representations, resulting in the overhead of constructing these models when applying the techniques to…

Cited by 128SourceScholar
2019

Towards Generalizing Sensorimotor Control Across Weather Conditions

IROS 2019poster

The ability of deep learning models to generalize well across different scenarios depends primarily on the quality and quantity of annotated data. Labeling large amounts of data for all possible scenarios that a model may encounter would not be feasible; if even possible. We propose a framework to d…

Cited by 7SourceScholar
2018

Discrete-Continuous ADMM for Transductive Inference in Higher-Order MRFs

CVPR 2018poster

This paper introduces a novel algorithm for transductive inference in higher-order MRFs, where the unary energies are parameterized by a variable classifier. The considered task is posed as a joint optimization problem in the continuous classifier parameters and the discrete label variables. In cont…

Cited by 10SourcePDFScholar