← Search

Andre Araujo

12 accepted papers

2026

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

CVPR 2026

Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features are poorly aligned across different modalities. For instance, the feature embedding for an RGB image and its corresponding depth map of the same scene exhibit

Cited by 0SourcecodeScholar
2026

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

CVPR 2026

Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation and depth prediction. However, a fundamental capability that these models still struggle with is aligning dense patch r

Cited by 0SourcecodeScholar
2025

AlignDiff: Learning Physically-Grounded Camera Alignment via Diffusion

ICCV 2025poster

Accurate camera calibration is a fundamental task for 3D perception, especially when dealing with real-world, in-the-wild environments where complex optical distortions are common. Existing methods often rely on pre-rectified images or calibration patterns, which limits their applicability and flexi…

Cited by 0SourcePDFScholar
2025

TIPS: Text-Image Pretraining with Spatial awareness

ICLR 2025poster

While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense visio…

2025

Tuning the Frequencies: Robust Training for Sinusoidal Neural Networks

CVPR 2025highlight

Sinusoidal neural networks have been shown effective as implicit neural representations (INRs) of low-dimensional signals, due to their smoothness and high representation capacity. However, initializing and training them remain empirical tasks which lack on deeper understanding to guide the learning…

Cited by 0SourcePDFScholar
2025

VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models

NeurIPS 2025poster

Foundation models have advanced computer vision by enabling strong performance across diverse tasks through large-scale pretraining and supervised fine-tuning. However, they may underperform in domains with distribution shifts and scarce labels, where supervised fine-tuning may be infeasible. While…

Cited by 0SourcecodeScholar
2024

UDON: Universal Dynamic Online distillatioN for generic image representations

NeurIPS 2024poster

Universal image representations are critical in enabling real-world fine-grained and instance-level recognition applications, where objects and entities from any domain must be identified at large scale. Despite recent advances, existing methods fail to capture important domain-specific knowledge, w…

2023

NAVI: Category-Agnostic Image Collections with High-Quality 3D Shape and Pose Annotations

NeurIPS 2023poster

Recent advances in neural reconstruction enable high-quality 3D object reconstruction from casually captured image collections. Current techniques mostly analyze their progress on relatively simple image collections where SfM techniques can provide ground-truth (GT) camera poses. We note that SfM te…

2023

Yes, we CANN: Constrained Approximate Nearest Neighbors for Local Feature-Based Visual Localization

ICCV 2023poster

Large-scale visual localization systems continue to relyon 3D point clouds built from image collections usingstructure-from-motion. While the 3D points in these modelsare represented using local image features, directly match-ing a query image's local features against the point cloud ischallenging d…

Cited by 8PDFcodeScholar
2020

Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval

CVPR 2020oral

While image retrieval and instance recognition techniques are progressing rapidly, there is a need for challenging datasets to accurately measure their performance -- while posing novel challenges that are relevant for practical applications. We introduce the Google Landmarks Dataset v2 (GLDv2), a n…

Cited by 438PDFcodeScholar
2019

Detect-To-Retrieve: Efficient Regional Aggregation for Image Search

CVPR 2019poster

Retrieving object instances among cluttered scenes efficiently requires compact yet comprehensive regional image representations. Intuitively, object semantics can help build the index that focuses on the most relevant regions. However, due to the lack of bounding-box datasets for objects of interes…

Cited by 157PDFcodeScholar
2017

Large-Scale Image Retrieval With Attentive Deep Local Features

ICCV 2017poster

We propose an attentive local feature descriptor suitable for large-scale image retrieval, referred to as DELF (DEep Local Feature). The new feature is based on convolutional neural networks, which are trained only with image-level annotations on a landmark image dataset. To identify semantically us…

Cited by 860PDFcodeScholar