← Search

Francesc Moreno-Noguer

45 accepted papers

2026

BEA-GS: BEyond RAdiance Supervision in 3DGS for Precise Object Extraction

CVPR 2026

Most Gaussian Splatting techniques that provide a 3D semantic representation of the scene don't optimize the underlying 3D geometry of the scene. This makes object-level editing or asset extraction challenging. Recent methods, like COBGS, Trace3D, and ObjectGS, acknowledge this limitation and propos

Cited by 0SourceScholar
2026

PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories

CVPR 2026

Recent advances in generative AI have dramatically improved photorealistic image synthesis, yet they fall short for studio-level multi-object compositing. This task demands simultaneous (i) near-perfect preservation of each item's identity, (ii) precise background and color fidelity, (iii) layout an

Cited by 0SourceScholar
2025

MEGA: Masked Generative Autoencoder for Human Mesh Recovery

CVPR 2025poster

Human Mesh Recovery (HMR) from a single RGB image is a highly ambiguous problem, as an infinite set of 3D interpretations can explain the 2D observation equally well. Nevertheless, most HMR methods overlook this issue and make a single prediction without accounting for this ambiguity. A few approach…

Cited by 1SourcePDFScholar
2024

"PoseEmbroider: Towards a 3D, Visual, Semantic-aware Human Pose Representation"

ECCV 2024poster

"Aligning multiple modalities in a latent space, such as images and texts, has shown to produce powerful semantic visual representations, fueling tasks like image captioning, text-to-image generation, or image grounding. In the context of human-centric vision, albeit CLIP-like representations encode…

Cited by 1SourcePDFScholar
2024

Estimating 3D Uncertainty Field: Quantifying Uncertainty for Neural Radiance Fields

ICRA 2024poster

Current methods based on Neural Radiance Fields (NeRF) significantly lack the capacity to quantify uncertainty in their predictions, particularly on the unseen space including the occluded and outside scene content. This limitation hinders their extensive applications in robotics, where the reliabil…

Cited by 11SourceScholar
2024

IReNe: Instant Recoloring of Neural Radiance Fields

CVPR 2024poster

Advances in NERFs have allowed for 3D scene reconstructions and novel view synthesis. Yet efficiently editing these representations while retaining photorealism is an emerging challenge. Recent methods face three primary limitations: they're slow for interactive use lack precision at object boundari…

Cited by 1SourcePDFScholar
2024

MultiPhys: Multi-Person Physics-aware 3D Motion Estimation

CVPR 2024poster

We introduce MultiPhys a method designed for recovering multi-person motion from monocular videos. Our focus lies in capturing coherent spatial placement between pairs of individuals across varying degrees of engagement. MultiPhys being physically aware exhibits robustness to jittering and occlusion…

Cited by 5SourcePDFScholar
2023

NeRFLight: Fast and Light Neural Radiance Fields Using a Shared Feature Grid

CVPR 2023poster

While original Neural Radiance Fields (NeRF) have shown impressive results in modeling the appearance of a scene with compact MLP architectures, they are not able to achieve real-time rendering. This has been recently addressed by either baking the outputs of NeRF into a data structure or arranging…

Cited by 6SourcePDFScholar
2023

PoseFix: Correcting 3D Human Poses with Natural Language

ICCV 2023poster

Automatically producing instructions to modify one's posture could open the door to endless applications, such as personalized coaching and in-home physical therapy. Tackling the reverse problem (i.e., refining a 3D pose based on some natural language feedback) could help for assisted 3D character a…

Cited by 29PDFScholar
2022

An Adaptable Approach to Learn Realistic Legged Locomotion without Examples

ICRA 2022poster

Learning controllers that reproduce legged locomotion in nature has been a longtime goal in robotics and computer graphics. While yielding promising results, recent approaches are not yet flexible enough to be applicable to legged systems of different morphologies. This is partly because they often…

Cited by 11SourceScholar
2022

Belief Revision Based Caption Re-ranker with Visual Semantic Information

COLING 2022main

In this work, we focus on improving the captions generated by image-caption generation systems. We propose a novel re-ranking approach that leverages visual-semantic measures to identify the ideal caption that maximally captures the visual information in the image. Our re-ranker utilizes the Belief…

2022

Conditional-Flow NeRF: Accurate 3D Modelling with Reliable Uncertainty Quantification

ECCV 2022poster

"A critical limitation of current methods based on Neural Radiance Fields (NeRF) is that they are unable to quantify the uncertainty associated with the learned appearance and geometry of the scene. This information is paramount in real applications such as medical diagnosis or autonomous driving wh…

2022

Context and Intention aware 3D Human Body Motion Prediction using an Attention Deep Learning model in Handover Tasks

IROS 2022poster

This work explores how contextual information and human intention affect the motion prediction of humans during a handover operation with a social robot. By classifying human intention in four different classes, we developed a model able to generate a different motion for each intention class. Furth…

Cited by 8SourceScholar
2022

LISA: Learning Implicit Shape and Appearance of Hands

CVPR 2022poster

This paper proposes a do-it-all neural model of human hands, named LISA. The model can capture accurate hand shape and appearance, generalize to arbitrary hand subjects, provide dense surface correspondences, be reconstructed from images in the wild and easily animated. We train LISA by minimizing t…

Cited by 80PDFScholar
2022

Learned Vertex Descent: A New Direction for 3D Human Model Fitting

ECCV 2022poster

"We propose a novel optimization-based paradigm for 3D human shape fitting on images. In contrast to existing approaches that directly regress the parameters of a low-dimensional statistical body model (e.g. SMPL) from input images, we propose training a deep network that, given solely image feature…

Cited by 40SourcePDFScholar
2022

PoseScript: 3D Human Poses from Natural Language

ECCV 2022poster

"Natural language is leveraged in many computer vision tasks such as image captioning, cross-modal retrieval or visual question answering, to provide fine-grained semantic information. While human pose is key to human understanding, current 3D human pose datasets lack detailed language descriptions.…

Cited by 66SourcePDFScholar
2022

Recognizing object surface material from impact sounds for robot manipulation

IROS 2022poster

We investigated the use of impact sounds generated during exploratory behaviors in a robotic manipulation setup as cues for predicting object surface material and for recognizing individual objects. We collected and make available the YCB-impact sounds dataset which includes over 3,000 impact sounds…

Cited by 12SourceScholar
2021

D-NeRF: Neural Radiance Fields for Dynamic Scenes

CVPR 2021poster

Neural rendering techniques combining machine learning with geometric reasoning have arisen as one of the most promising approaches for synthesizing novel views of a scene from a sparse set of images. Among these, stands out the Neural radiance fields (NeRF), which trains a deep network to map 5D in…

Cited by 1585PDFScholar
2021

Generating Attribution Maps With Disentangled Masked Backpropagation

ICCV 2021poster

Attribution map visualization has arisen as one of the most effective techniques to understand the underlying inference process of Convolutional Neural Networks. In this task, the goal is to compute an score for each image pixel related to its contribution to the network output. In this paper, we in…

Cited by 4PDFcodeScholar
2021

H3D-Net: Few-Shot High-Fidelity 3D Head Reconstruction

ICCV 2021poster

Recent learning approaches that implicitly represent surface geometry using coordinate-based neural representations have shown impressive results in the problem of multi-view 3D reconstruction. The effectiveness of these techniques is, however, subject to the availability of a large number (several…

Cited by 104PDFScholar
2021

Multi-FinGAN: Generative Coarse-To-Fine Sampling of Multi-Finger Grasps

ICRA 2021poster

While there exists many methods for manipulating rigid objects with parallel-jaw grippers, grasping with multi-finger robotic hands remains a quite unexplored research topic. Reasoning and planning collision-free trajectories on the additional degrees of freedom of several fingers represents an impo…

Cited by 63SourcecodeScholar
2021

SMPLicit: Topology-Aware Generative Model for Clothed People

CVPR 2021poster

In this paper we introduce SMPLicit, a novel generative model to jointly represent body pose, shape and clothing geometry. In contrast to existing learning-based approaches that require training specific models for each type of garment, SMPLicit can represent in a unified manner different garment to…

Cited by 219PDFcodeScholar
2021

Uncertainty-Aware Camera Pose Estimation From Points and Lines

CVPR 2021poster

Perspective-n-Point-and-Line (PnPL) algorithms aim at fast, accurate, and robust camera localization with respect to a 3D model from 2D-3D feature correspondences, being a major part of modern robotic and AR/VR systems. Current point-based pose estimation methods use only 2D feature detection uncert…

Cited by 30PDFcodeScholar
2020

3D Human Shape and Pose from a Single Low-Resolution Image with Self-Supervised Learning

ECCV 2020poster

3D human shape and pose estimation from monocular images has been an active area of research in computer vision, having a substantial impact on the development of new applications, from activity recognition to creating virtual avatars. Existing deep learning methods for 3D human shape and pose estim…

2020

C-Flow: Conditional Generative Flow Models for Images and 3D Point Clouds

CVPR 2020poster

Flow-based generative models have highly desirable properties like exact log-likelihood evaluation and exact latent-variable inference, however they are still in their infancy and have not received as much attention as alternative generative models. In this paper, we introduce C-Flow, a novel condit…

Cited by 113PDFScholar
2020

GanHand: Predicting Human Grasp Affordances in Multi-Object Scenes

CVPR 2020oral

The rise of deep learning has brought remarkable progress in estimating hand geometry from images where the hands are part of the scene. This paper focuses on a new problem not explored so far, consisting in predicting how a human would grasp one or several objects, given a single RGB image of these…

Cited by 201PDFScholar
2019

3DPeople: Modeling the Geometry of Dressed Humans

ICCV 2019poster

Recent advances in 3D human shape estimation build upon parametric representations that model very well the shape of the naked body, but are not appropriate to represent the clothing geometry. In this paper, we present an approach to model dressed humans and predict their geometry from single images…

Cited by 152PDFScholar
2018

GANimation: Anatomically-aware Facial Animation from a Single Image

ECCV 2018poster

Recent advances in Generative Adversarial Networks (GANs) have shown impressive results for task of facial expression synthesis. The most successful architecture is StarGAN, that conditions GANs' generation process with images of a specific domain, namely a set of images of persons sharing the same…

2018

Geometry-Aware Network for Non-Rigid Shape Prediction From a Single View

CVPR 2018poster

We propose a method for predicting the 3D shape of a deformable surface from a single view. By contrast with previous approaches, we do not need a pre-registered template of the surface, and our method is robust to the lack of texture and partial occlusions. At the core of our approach is a geometry…

Cited by 66SourcePDFScholar
2018

Image Collection Pop-Up: 3D Reconstruction and Clustering of Rigid and Non-Rigid Categories

CVPR 2018poster

This paper introduces an approach to simultaneously estimate 3D shape, camera pose, and object and type of deformation clustering, from partial 2D annotations in a multi-instance collection of images. Furthermore, we can indistinctly process rigid and non-rigid categories. This advances existing wor…

Cited by 31SourcePDFScholar
2018

Unsupervised Person Image Synthesis in Arbitrary Poses

CVPR 2018poster

We present a novel approach for synthesizing photo-realistic images of people in arbitrary poses using generative adversarial learning. Given an input image of a person and a desired pose represented by a 2D skeleton, our model renders the image of the same person under the new pose, synthesizing no…

Cited by 211SourcePDFScholar
2017

DUST: Dual Union of Spatio-Temporal Subspaces for Monocular Multiple Object 3D Reconstruction

CVPR 2017poster

We present an approach to reconstruct the 3D shape of multiple deforming objects from incomplete 2D trajectories acquired by a single camera. Additionally, we simultaneously provide spatial segmentation (i.e., we identify each of the objects in every frame) and temporal clustering (i.e., we split th…

Cited by 37PDFScholar
2017

Depth-aware convolutional neural networks for accurate 3D pose estimation in RGB-D images

IROS 2017poster

Most recent approaches to 3D pose estimation from RGB-D images address the problem in a two-stage pipeline. First, they learn a classifier-typically a random forest-to predict the position of each input pixel on the object surface. These estimates are then used to define an energy function that is m…

Cited by 16SourceScholar
2017

Learning Depth-Aware Deep Representations for Robotic Perception

RA-L 2017

Exploiting RGB-D data by means of convolutional neural networks (CNNs) is at the core of a number of robotics applications, including object detection, scene semantic segmentation, and grasping. Most existing approaches, however, exploit RGB-D data by simply considering depth as an additional input

Cited by 33SourceScholar
2015

Discriminative Learning of Deep Convolutional Feature Point Descriptors

ICCV 2015poster

Deep learning has revolutionalized image-level tasks such as classification, but patch-level tasks, such as correspondence, still rely on hand-crafted features, e.g. SIFT. In this paper we use Convolutional Neural Networks (CNNs) to learn discriminant patch representations and in particular train a…

Cited by 1022PDFcodeScholar
2015

Neuroaesthetics in Fashion: Modeling the Perception of Fashionability

CVPR 2015poster

In this paper, we analyze the fashion of clothing of a large social website. Our goal is to learn and predict how fashionable a person looks on a photograph and suggest subtle improvements the user could make to improve her/his appeal. We propose a Conditional Random Field model that jointly reasons…

Cited by 251SourcePDFScholar