← Search

Silvio Giancola

13 accepted papers

2025

3D Convex Splatting: Radiance Field Rendering with 3D Smooth Convexes

CVPR 2025highlight

Recent advances in radiance field reconstruction, such as 3D Gaussian Splatting (3DGS), have achieved high-quality novel view synthesis and fast rendering by representing scenes with compositions of Gaussian primitives. However, 3D Gaussians present several limitations for scene reconstruction. Accu…

2024

Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders

ECCV 2024poster

"Self-supervised pre-training of image encoders is omnipresent in the literature, particularly following the introduction of Masked autoencoders (MAE). Current efforts attempt to learn object-centric representations from motion in videos. In particular, SiamMAE recently introduced a Siamese network,…

2024

TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks

ECCV 2024poster

"Neural radiance fields (NeRFs) generally require many images with accurate poses for accurate novel view synthesis, which does not reflect realistic setups where views can be sparse and poses can be noisy. Previous solutions for learning NeRFs with sparse views and noisy poses only consider local g…

2023

EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual Queries

ICCV 2023oral

With the recent advances in video and 3D understanding, novel 4D spatio-temporal methods fusing both concepts have emerged. Towards this direction, the Ego4D Episodic Memory Benchmark proposed a task for Visual Queries with 3D Localization (VQ3D). Given an egocentric video clip and an image crop dep…

Cited by 21PDFcodeScholar
2023

Voint Cloud: Multi-View Point Cloud Representation for 3D Understanding

ICLR 2023poster

Multi-view projection methods have demonstrated promising performance on 3D understanding tasks like 3D classification and segmentation. However, it remains unclear how to combine such multi-view methods with the widely available 3D point clouds. Previous methods use unlearned heuristics to combine…

2022

3DeformRS: Certifying Spatial Deformations on Point Clouds

CVPR 2022poster

3D computer vision models are commonly used in security-critical applications such as autonomous driving and surgical robotics. Emerging concerns over the robustness of these models against real-world deformations must be addressed practically and reliably. In this work, we propose 3DeformRS, a meth…

Cited by 14PDFcodeScholar
2022

MAD: A Scalable Dataset for Language Grounding in Videos From Movie Audio Descriptions

CVPR 2022poster

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of these datasets for the video-language grounding task. Recen…

Cited by 123PDFcodeScholar
2022

Real-Time Hyperspectral Imaging in Hardware via Trained Metasurface Encoders

CVPR 2022poster

Hyperspectral imaging has attracted significant attention to identify spectral signatures for image classification and automated pattern recognition in computer vision. State-of-the-art implementations of snapshot hyperspectral imaging rely on bulky, non-integrated, and expensive optical elements, i…

Cited by 30PDFcodeScholar
2022

SCTN: Sparse Convolution-Transformer Network for Scene Flow Estimation

AAAI 2022technical

We propose a novel scene flow estimation approach to capture and infer 3D motions from point clouds. Estimating 3D motions for point clouds is challenging, since a point cloud is unordered and its density is significantly non-uniform. Such unstructured data poses difficulties in matching correspondi…

2020

A Context-Aware Loss Function for Action Spotting in Soccer Videos

CVPR 2020poster

In video understanding, action spotting consists in temporally localizing human-induced events annotated with single timestamps. In this paper, we propose a novel loss function that specifically considers the temporal context naturally present around each action, rather than focusing on the single a…

Cited by 111PDFcodeScholar
2018

TrackingNet: A Large-Scale Dataset and Benchmark for Object Tracking in the Wild

ECCV 2018poster

Despite the numerous developments in object tracking, further development of current tracking algorithms is limited by small and mostly saturated datasets. As a matter of fact, data-hungry trackers based on deep-learning currently rely on object detection datasets due to the scarcity of dedicated la…

Cited by 1191SourcePDFScholar