← Search

Orazio Gallo

17 accepted papers

2026

4DP-QA: Scalable QA for 4D Perception in Vision Language Models

CVPR 2026

Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself, is further complicated by two factors. First, VLMs observe motion indirectly via its projection onto 2D images. Second

Cited by 0SourceScholar
2025

FoundationStereo: Zero-Shot Stereo Matching

CVPR 2025award

Tremendous progress has been made in deep stereo matching to excel on benchmark datasets through per-domain fine-tuning. However, achieving strong zero-shot generalization - a hallmark of foundation models in other computer vision tasks - remains challenging for stereo matching. We introduce Foundat…

2025

Unmasking Puppeteers: Leveraging Biometric Leakage to Expose Impersonation in AI-Based Videoconferencing

NeurIPS 2025poster

AI-based talking-head videoconferencing systems reduce bandwidth by transmitting a latent representation of a speaker’s pose and expression, which is used to synthesize frames on the receiver's end. However, these systems are vulnerable to “puppeteering” attacks, where an adversary controls the iden…

Cited by 0SourceScholar
2025

Zero-Shot Monocular Scene Flow Estimation in the Wild

CVPR 2025award

Large models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow.Even though scene flow prediction has wide potential, its practical use is limited because of the lack of generalization of current predictiv…

Cited by 1SourcePDFScholar
2024

Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos

ECCV 2024poster

"Modern avatar generators allow anyone to synthesize photorealistic real-time talking avatars, ushering in a new era of avatar-based human communication, such as with immersive AR/VR interactions or videoconferencing with limited bandwidths. Their safe adoption, however, requires a mechanism to veri…

Cited by 3SourcePDFScholar
2023

Zero-Shot Pose Transfer for Unrigged Stylized 3D Characters

CVPR 2023poster

Transferring the pose of a reference avatar to stylized 3D characters of various shapes is a fundamental task in computer graphics. Existing methods either require the stylized characters to be rigged, or they use the stylized character in the desired pose as ground truth at training. We present a z…

2022

Efficient Geometry-Aware 3D Generative Adversarial Networks

CVPR 2022oral

Unsupervised generation of high-quality multi-view-consistent images and 3D shapes using only collections of single-view 2D photographs has been a long-standing challenge. Existing 3D GANs are either compute-intensive or make approximations that are not 3D-consistent; the former limits quality and r…

Cited by 1564PDFcodeScholar
2022

Watch It Move: Unsupervised Discovery of 3D Joints for Re-Posing of Articulated Objects

CVPR 2022poster

Rendering articulated objects while controlling their poses is critical to applications such as virtual reality or animation for movies. Manipulating the pose of an object, however, requires the understanding of its underlying structure, that is, its joints and how they interact with each other. Unf…

Cited by 51PDFcodeScholar
2020

Bi3D: Stereo Depth Estimation via Binary Classifications

CVPR 2020poster

Stereo-based depth estimation is a cornerstone of computer vision, with state-of-the-art methods delivering accurate results in real time. For several applications such as autonomous navigation, however, it may be useful to trade accuracy for lower latency. We present Bi3D, a method that estimates d…

Cited by 103PDFcodeScholar
2020

Generative View Synthesis: From Single-view Semantics to Novel-view Images

NeurIPS 2020poster

Content creation, central to applications such as virtual reality, can be tedious and time-consuming. Recent image synthesis methods simplify this task by offering tools to generate new views from as little as a single input image, or by converting a semantic map into a photorealistic image. We pro…

2020

Novel View Synthesis of Dynamic Scenes With Globally Coherent Depths From a Monocular Camera

CVPR 2020poster

This paper presents a new method to synthesize an image from arbitrary views and times given a collection of images of a dynamic scene. A key challenge for the novel view synthesis arises from dynamic scene reconstruction where epipolar geometry does not apply to the local motion of dynamic contents…

Cited by 172PDFScholar
2019

Pixel-Adaptive Convolutional Neural Networks

CVPR 2019poster

Convolutions are the fundamental building blocks of CNNs. The fact that their weights are spatially shared is one of the main reasons for their widespread use, but it is also a major limitation, as it makes convolutions content-agnostic. We propose a pixel-adaptive convolution (PAC) operation, a sim…

Cited by 383PDFcodeScholar
2018

Separating Reflection and Transmission Images in the Wild

ECCV 2018poster

The reflections caused by common semi-reflectors, such as glass windows, can impact the performance of computer vision algorithms. State-of-the-art methods can remove reflections on synthetic data and in controlled scenarios. However, they are based on strong assumptions and do not generalize well t…

Cited by 80SourcePDFScholar
2018

Tackling 3D ToF Artifacts Through Learning and the FLAT Dataset

ECCV 2018poster

Scene motion, multiple reflections, and sensor noise introduce artifacts in the depth reconstruction performed by time-of-flight cameras. We propose a two-stage, deep-learning approach to address all of these sources of artifacts simultaneously. We also introduce FLAT, a synthetic dataset of 2000 To…

Cited by 66SourcePDFScholar