← Search

Paolo Rota

11 accepted papers

2026

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

CVPR 2026

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this problem, we introduce TerraScope, a unified VLM that delivers pixel-grounded geospa

Cited by 0SourceScholar
2025

ConViS-Bench: Estimating Video Similarity Through Semantic Concepts

NeurIPS 2025poster

What does it mean for two videos to be similar? Videos may appear similar when judged by the actions they depict, yet entirely different if evaluated based on the locations where they were filmed. While humans naturally compare videos by taking different aspects into account, this ability has not be…

Cited by 0SourceScholar
2025

ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression

NeurIPS 2025oral

The hypothesis that Convolutional Neural Networks (CNNs) are inherently texture-biased has shaped much of the discourse on feature use in deep learning. We revisit this hypothesis by examining limitations in the cue-conflict experiment by Geirhos et al. To address these limitations, we propose a dom…

Cited by 0SourcecodeScholar
2025

Multi-focal Conditioned Latent Diffusion for Person Image Synthesis

CVPR 2025poster

The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, the compression process of LDM often results in the deterioration of details, pa…

2025

On Large Multimodal Models as Open-World Image Classifiers

ICCV 2025poster

Traditional image classification requires a predefined list of semantic categories. In contrast, Large Multimodal Models (LMMs) can sidestep this requirement by classifying images directly using natural language (e.g., answering the prompt "What is the main object in the image?"). Despite this remar…

2024

Test-Time Zero-Shot Temporal Action Localization

CVPR 2024poster

Zero-Shot Temporal Action Localization (ZS-TAL) seeks to identify and locate actions in untrimmed videos unseen during training. Existing ZS-TAL methods involve fine-tuning a model on a large amount of annotated training data. While effective training-based ZS-TAL approaches assume the availability…

2023

AutoLabel: CLIP-Based Framework for Open-Set Video Domain Adaptation

CVPR 2023poster

Open-set Unsupervised Video Domain Adaptation (OUVDA) deals with the task of adapting an action recognition model from a labelled source domain to an unlabelled target domain that contains "target-private" categories, which are present in the target but absent in the source. In this work we deviate…

2023

Rotation Synchronization via Deep Matrix Factorization

ICRA 2023poster

In this paper we address the rotation synchronization problem, where the objective is to recover absolute rotations starting from pairwise ones, where the unknowns and the measures are represented as nodes and edges of a graph, respectively. This problem is an essential task for structure from motio…

Cited by 12SourcecodeScholar
2023

The Unreasonable Effectiveness of Large Language-Vision Models for Source-Free Video Domain Adaptation

ICCV 2023poster

Source-Free Video Unsupervised Domain Adaptation (SFVUDA) task consists in adapting an action recognition model, trained on a labelled source dataset, to an unlabelled target dataset, without accessing the actual source data. The previous approaches have attempted to address SFVUDA by leveraging sel…

Cited by 11PDFcodeScholar
2023

Vocabulary-free Image Classification

NeurIPS 2023poster

Recent advances in large vision-language models have revolutionized the image classification paradigm. Despite showing impressive zero-shot capabilities, a pre-defined set of categories, a.k.a. the vocabulary, is assumed at test time for composing the textual prompts. However, such assumption can be…

2015

The S-Hock Dataset: Analyzing Crowds at the Stadium

CVPR 2015poster

The topic of crowd modeling in computer vision usually assumes a single generic typology of crowd, which is very simplistic. In this paper we adopt a taxonomy that is widely accepted in sociology, focusing on a particular category, the spectator crowd, which is formed by people "interested in watchi…

Cited by 57SourcePDFScholar