← Search

Gabriele Berton

10 accepted papers

2026

A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

CVPR 2026

Anticipating diverse future states is a central challenge in video world modeling. Discriminative world models produce deterministic predictions that implicitly average over possible futures, while existing generative world models remain computationally expensive. Recent work demonstrates that predi

Cited by 0SourcecodeScholar
2026

AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localization

CVPR 2026

Precise and real-time visual localization is critical for applications like AR/VR and robotics, especially on resource-constrained edge devices such as smart glasses, where battery life and heat dissipation can be primary concerns. While many efficient models exist, further reducing compute without

Cited by 0SourceScholar
2024

EarthLoc: Astronaut Photography Localization by Indexing Earth from Space

CVPR 2024poster

Astronaut photography spanning six decades of human spaceflight presents a unique Earth observations dataset with immense value for both scientific research and disaster response. Despite its significance accurately localizing the geographical extent of these images crucial for effective utilization…

2024

Scale-Free Image Keypoints Using Differentiable Persistent Homology

ICML 2024poster

In computer vision, keypoint detection is a fundamental task, with applications spanning from robotics to image retrieval; however, existing learning-based methods suffer from scale dependency, and lack flexibility. This paper introduces a novel approach that leverages Morse theory and persistent ho…

2023

Divide&Classify: Fine-Grained Classification for City-Wide Visual Geo-Localization

ICCV 2023poster

Visual Place recognition is commonly addressed as an image retrieval problem. However, retrieval methods are impractical to scale to large datasets, densely sampled from city-wide maps, since their dimension impact negatively on the inference time. Using approximate nearest neighbour search for retr…

Cited by 14PDFcodeScholar
2023

EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition

ICCV 2023poster

Visual Place Recognition is a task that aims to predict the place of an image (called query) based solely on its visual features. This is typically done through image retrieval, where the query is matched to the most similar images from a large database of geotagged photos, using learned global des…

Cited by 87PDFcodeScholar
2022

Deep Visual Geo-Localization Benchmark

CVPR 2022oral

In this paper, we propose a new open-source benchmarking framework for Visual Geo-localization (VG) that allows to build, train, and test a wide range of commonly used architectures, with the flexibility to change individual components of a geo-localization pipeline. The purpose of this framework is…

Cited by 103PDFcodeScholar
2021

Viewpoint Invariant Dense Matching for Visual Geolocalization

ICCV 2021poster

In this paper we propose a novel method for image matching based on dense local features and tailored for visual geolocalization. Dense local features matching is robust against changes in illumination and occlusions, but not against viewpoint shifts which are a fundamental aspect of geolocalization…

Cited by 37PDFcodeScholar