← Search

João F. Henriques

16 accepted papers

2025

TrafficLoc: Localizing Traffic Surveillance Cameras in 3D Scenes

ICCV 2025poster

We tackle the problem of localizing traffic cameras within a 3D reference map and propose a novel image-to-point cloud registration (I2P) method, TrafficLoc, in a coarse-to-fine matching fashion. To overcome the lack of large-scale real-world intersection datasets, we first introduce Carla Intersect…

2024

A Sound Approach: Using Large Language Models to Generate Audio Descriptions for Egocentric Text-Audio Retrieval

ICASSP 2024accepted

Video databases from the internet are a valuable source of text-audio retrieval datasets. However, given that sound and vision streams represent different "views" of the data, treating visual descriptions as audio descriptions is far from optimal. Even if audio class labels are present, they commonl…

Cited by 0SourceScholar
2023

A Light Touch Approach to Teaching Transformers Multi-View Geometry

CVPR 2023poster

Transformers are powerful visual learners, in large part due to their conspicuous lack of manually-specified priors. This flexibility can be problematic in tasks that involve multiple-view geometry, due to the near-infinite possible variations in 3D shapes and viewpoints (requiring flexibility), and…

Cited by 8SourcePDFScholar
2023

CASSPR: Cross Attention Single Scan Place Recognition

ICCV 2023poster

Place recognition based on point clouds (LiDAR) is an important component for autonomous robots or self-driving vehicles. Current SOTA performance is achieved on accumulated LiDAR submaps using either point-based or voxel-based structures. While voxel-based approaches nicely integrate spatial contex…

Cited by 68PDFcodeScholar
2022

SNeS: Learning Probably Symmetric Neural Surfaces from Incomplete Data

ECCV 2022poster

"We present a method for the accurate 3D reconstruction of partly-symmetric objects. We build on the strengths of recent advances in neural reconstruction and rendering such as Neural Radiance Fields (NeRF). A major shortcomings of such approaches is that they fail to reconstruct any part of the obj…

2021

On Compositions of Transformations in Contrastive Self-Supervised Learning

ICCV 2021poster

In the image domain, excellent representations can be learned by inducing invariance to content-preserving transformations via noise contrastive learning. In this paper, we generalize contrastive learning to a wider set of transformations, and their compositions, for which either invariance or disti…

Cited by 73PDFcodeScholar
2021

QUERYD: A Video Dataset with High-Quality Text and Audio Narrations

ICASSP 2021accepted

We introduce QuerYD, a new large-scale dataset for retrieval and event localisation in video. A unique feature of our dataset is the availability of two audio tracks for each video: the original audio, and a high-quality spoken description of the visual content. The dataset is based on YouDescribe […

Cited by 0SourceScholar
2021

Space-Time Crop & Attend: Improving Cross-Modal Video Representation Learning

ICCV 2021poster

The quality of the image representations obtained from self-supervised learning depends strongly on the type of data augmentations used in the learning formulation. Recent papers have ported these methods from still images to videos and found that leveraging both audio and video signals yields stron…

Cited by 43PDFcodeScholar
2016

Learning feed-forward one-shot learners

NeurIPS 2016poster

One-shot learning is usually tackled by using generative models or discriminative embeddings. Discriminative methods based on deep learning, which are very effective in other learning scenarios, are ill-suited for one-shot learning as they need large amounts of training data. In this paper, we propo…

Cited by 576SourcePDFScholar