← Search

Guillermo Garcia-Hernando

13 accepted papers

2026

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

CVPR 2026

We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level and by satellite. Faithful reconstructions require good camera coverage, but ground imagery is time-consuming and hard to capture at scale for large outdoor s

Cited by 0SourcecodeScholar
2025

MVSAnywhere: Zero-Shot Multi-View Stereo

CVPR 2025poster

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision.However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs outdoor). Training a general-purpose multi-view stereo model is challenging an…

2025

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

ICCV 2025poster

We introduce the task of Language-Guided Object Placement in Real 3D Scenes. Given a 3D reconstructed point-cloud scene, a 3D asset, and a natural-language instruction, the goal is to place the asset so that the instruction is satisfied. The task demands tackling four intertwined challenges: (a) one…

Cited by 0SourcePDFScholar
2024

DoubleTake: Geometry Guided Depth Estimation

ECCV 2024poster

"Estimating depth from a sequence of posed RGB images is a fundamental computer vision task, with applications in augmented reality, path planning etc. Prior work typically makes use of previous frames in a multi view stereo framework, relying on matching textures in a local neighborhood. In contras…

Cited by 1SourcePDFScholar
2023

Removing Objects From Neural Radiance Fields

CVPR 2023poster

Neural Radiance Fields (NeRFs) are emerging as a ubiquitous scene representation that allows for novel view synthesis. Increasingly, NeRFs will be shareable with other people. Before sharing a NeRF, though, it might be desirable to remove personal information or unsightly objects. Such removal is no…

Cited by 71SourcePDFScholar
2022

Map-Free Visual Relocalization: Metric Pose Relative to a Single Image

ECCV 2022poster

"Can we relocalize in a scene represented by a single reference image? Standard visual relocalization requires hundreds of images and scale calibration to build a scene-specific 3D map. In contrast, we propose Map-free Relocalization, i.e., using only one photo of a scene to enable instant, metric s…

2020

Active 6D Multi-Object Pose Estimation in Cluttered Scenarios with Deep Reinforcement Learning

IROS 2020

In this work, we explore how a strategic selection of camera movements can facilitate the task of 6D multi-object pose estimation in cluttered scenarios while respecting real-world constraints such as time and distance travelled, important in robotics and augmented reality applications. In the propo

Cited by 12SourceScholar
2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement Learning

IROS 2020poster

Dexterous manipulation of objects in virtual environments with our bare hands, by using only a depth sensor and a state-of-the-art 3D hand pose estimator (HPE), is challenging. While virtual environments are ruled by physics, e.g. object weights and surface frictions, the absence of force feedback m…

Cited by 62SourceScholar
2020

Predicting Visual Overlap of Images Through Interpretable Non-Metric Box Embeddings

ECCV 2020poster

To what extent are two images picturing the same 3D surfaces? Even when this is a known scene, the answer typically requires an expensive search across scale space, with matching and geometric verification of large sets of local features. This expense is further multiplied when a query image is eval…

2018

Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals

CVPR 2018poster

In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods o…

Cited by 277SourcePDFScholar
2018

First-Person Hand Action Benchmark With RGB-D Videos and 3D Hand Pose Annotations

CVPR 2018poster

In this work we study the use of 3D hand poses to recognize first-person dynamic hand actions interacting with 3D objects. Towards this goal, we collected RGB-D video sequences comprised of more than 100K frames of 45 daily hand action categories, involving 26 different objects in several hand conf…

Cited by 672SourcePDFScholar
2017

Transition Forests: Learning Discriminative Temporal Transitions for Action Recognition and Detection

CVPR 2017poster

A human action can be seen as transitions between one's body poses over time, where the transition depicts a temporal relation between two poses. Recognizing actions thus involves learning a classifier sensitive to these pose transitions as well as to static poses. In this paper, we introduce a nove…

Cited by 102PDFScholar