← Search

Antonio Agudo

23 accepted papers

2026

JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory Generation

ICLR 2026poster

Generative models often treat continuous data and discrete events as separate processes, creating a gap in modeling complex systems where they interact synchronously. To bridge this gap, we introduce $\textbf{JointDiff}$, a novel diffusion framework designed to unify these two processes by simultane…

Cited by 0SourcecodeScholar
2025

Dual-Space Augmented Intrinsic-LoRA for Wind Turbine Segmentation

ICASSP 2025accepted

Accurate segmentation of wind turbine blade (WTB) images is critical for effective assessments, as it directly influences the performance of automated damage detection systems. Despite advancements in large universal vision models, these models often underperform in domain-specific tasks like WTB se…

Cited by 0SourceScholar
2025

Unified Uncertainty-Aware Diffusion for Multi-Agent Trajectory Modeling

CVPR 2025poster

Multi-agent trajectory modeling has primarily focused on forecasting future states, often overlooking broader tasks like trajectory completion, which are crucial for real-world applications such as correcting tracking data. Existing methods also generally predict agents' states without offering any…

2024

VQ-HPS: Human Pose and Shape Estimation in a Vector-Quantized Latent Space

ECCV 2024poster

"Previous works on Human Pose and Shape Estimation (HPSE) from RGB images can be broadly categorized into two main groups: parametric and non-parametric approaches. Parametric techniques leverage a low-dimensional statistical body model for realistic results, whereas recent non-parametric methods ac…

2023

Detail-Aware Uncalibrated Photometric Stereo

ICASSP 2023accepted

Photometric stereo is the problem of jointly inferring the 3D reconstruction, reflectance, lighting and specularities of an object from a set of visual signals. Recently, some variational, uncalibrated, unsupervised and unified formulations have provided robust solutions to the problem while reducin…

Cited by 0SourceScholar
2023

On discrete symmetries of robotics systems: A group-theoretic and data-driven analysis

RSS 2023poster

We present a comprehensive study on discrete morphological symmetries of dynamical systems, which are commonly observed in biological and artificial locomoting systems, such as legged, swimming, and flying animals/robots/virtual characters. These symmetries arise from the presence of one or more pla…

2022

An Adaptable Approach to Learn Realistic Legged Locomotion without Examples

ICRA 2022poster

Learning controllers that reproduce legged locomotion in nature has been a longtime goal in robotics and computer graphics. While yielding promising results, recent approaches are not yet flexible enough to be applicable to legged systems of different morphologies. This is partly because they often…

Cited by 11SourceScholar
2022

Conditional-Flow NeRF: Accurate 3D Modelling with Reliable Uncertainty Quantification

ECCV 2022poster

"A critical limitation of current methods based on Neural Radiance Fields (NeRF) is that they are unable to quantify the uncertainty associated with the learned appearance and geometry of the scene. This information is paramount in real applications such as medical diagnosis or autonomous driving wh…

2021

Generating Attribution Maps With Disentangled Masked Backpropagation

ICCV 2021poster

Attribution map visualization has arisen as one of the most effective techniques to understand the underlying inference process of Convolutional Neural Networks. In this task, the goal is to compute an score for each image pixel related to its contribution to the network output. In this paper, we in…

Cited by 4PDFcodeScholar
2021

Uncertainty-Aware Camera Pose Estimation From Points and Lines

CVPR 2021poster

Perspective-n-Point-and-Line (PnPL) algorithms aim at fast, accurate, and robust camera localization with respect to a 3D model from 2D-3D feature correspondences, being a major part of modern robotic and AR/VR systems. Current point-based pose estimation methods use only 2D feature detection uncert…

Cited by 30PDFcodeScholar
2020

Neural Dense Non-Rigid Structure from Motion with Latent Space Constraints

ECCV 2020poster

We introduce the first dense neural non-rigid structure from motion (N-NRSfM) approach, which can be trained end-to-end in an unsupervised manner from 2D point tracks. Compared to the competing methods, our combination of loss functions is fully-differentiable and can be readily integrated into deep…

Cited by 66SourcePDFScholar
2018

GANimation: Anatomically-aware Facial Animation from a Single Image

ECCV 2018poster

Recent advances in Generative Adversarial Networks (GANs) have shown impressive results for task of facial expression synthesis. The most successful architecture is StarGAN, that conditions GANs' generation process with images of a specific domain, namely a set of images of persons sharing the same…

2018

Geometry-Aware Network for Non-Rigid Shape Prediction From a Single View

CVPR 2018poster

We propose a method for predicting the 3D shape of a deformable surface from a single view. By contrast with previous approaches, we do not need a pre-registered template of the surface, and our method is robust to the lack of texture and partial occlusions. At the core of our approach is a geometry…

Cited by 66SourcePDFScholar
2018

Image Collection Pop-Up: 3D Reconstruction and Clustering of Rigid and Non-Rigid Categories

CVPR 2018poster

This paper introduces an approach to simultaneously estimate 3D shape, camera pose, and object and type of deformation clustering, from partial 2D annotations in a multi-instance collection of images. Furthermore, we can indistinctly process rigid and non-rigid categories. This advances existing wor…

Cited by 31SourcePDFScholar
2018

Unsupervised Person Image Synthesis in Arbitrary Poses

CVPR 2018poster

We present a novel approach for synthesizing photo-realistic images of people in arbitrary poses using generative adversarial learning. Given an input image of a person and a desired pose represented by a 2D skeleton, our model renders the image of the same person under the new pose, synthesizing no…

Cited by 211SourcePDFScholar
2017

DUST: Dual Union of Spatio-Temporal Subspaces for Monocular Multiple Object 3D Reconstruction

CVPR 2017poster

We present an approach to reconstruct the 3D shape of multiple deforming objects from incomplete 2D trajectories acquired by a single camera. Additionally, we simultaneously provide spatial segmentation (i.e., we identify each of the objects in every frame) and temporal clustering (i.e., we split th…

Cited by 37PDFScholar
2017

PL-SLAM: Real-time monocular visual SLAM with points and lines

ICRA 2017poster

Low textured scenes are well known to be one of the main Achilles heels of geometric computer vision algorithms relying on point correspondences, and in particular for visual SLAM. Yet, there are many environments in which, despite being low textured, one can still reliably estimate line-based geome…

Cited by 551SourceScholar