← Search

Christos Sakaridis

25 accepted papers

2026

Robust Promptable Video Object Segmentation

CVPR 2026

The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in safety-critical domains. This paper offers the first comprehensive study on robust PVOS (RobustPVOS). We first construct a new, comprehensive benchm

Cited by 0SourcecodeScholar
2025

CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes

RA-L 2025

Leveraging multiple sensors is crucial for robust semantic perception in autonomous driving, as each sensor type has complementary strengths and weaknesses. However, existing sensor fusion methods often treat sensors uniformly across all conditions, leading to suboptimal performance. By contrast, we

Cited by 37SourcecodeScholar
2025

PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields

CVPR 2025poster

We tackle the ill-posed inverse rendering problem in 3D reconstruction with a Neural Radiance Field (NeRF) approach informed by Physics-Based Rendering (PBR) theory, named PBR-NeRF. Our method addresses a key limitation in most NeRF and 3D Gaussian Splatting approaches: they estimate view-dependent…

2025

UniK3D: Universal Camera Monocular 3D Estimation

CVPR 2025poster

Monocular 3D estimation is crucial for visual perception. However, current methods fall short by relying on oversimplified assumptions, such as pinhole camera models or rectified images. These limitations severely restrict their general applicability, causing poor performance in real-world scenarios…

2024

Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding

ECCV 2024poster

"3D visual grounding is the task of localizing the object in a 3D scene which is referred by a description in natural language. With a wide range of applications ranging from autonomous indoor robotics to AR/VR, the task has recently risen in popularity. A common formulation to tackle 3D visual grou…

2024

MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty

ECCV 2024poster

"Achieving level-5 driving automation in autonomous vehicles necessitates a robust semantic visual perception system capable of parsing data from different sensors across diverse conditions. However, existing semantic perception datasets often lack important non-camera modalities typically used in a…

2024

MuRF: Multi-Baseline Radiance Fields

CVPR 2024poster

We present Multi-Baseline Radiance Fields (MuRF) a general feed-forward approach to solving sparse view synthesis under multiple different baseline settings (small and large baselines and different number of input views). To render a target novel view we discretize the 3D space into planes parallel…

2024

UniDepth: Universal Monocular Metric Depth Estimation

CVPR 2024highlight

Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling. However the remarkable accuracy of recent MMDE methods is confined to their training domains. These methods fail to generalize to unseen domains even in the presence of moderate do…

2024

Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes

CVPR 2024highlight

The estimation of implicit cross-frame correspondences and the high computational cost have long been major challenges in video semantic segmentation (VSS) for driving scenes. Prior works utilize keyframes feature propagation or cross-frame attention to address these issues. By contrast we are the f…

2023

Contrastive Model Adaptation for Cross-Condition Robustness in Semantic Segmentation

ICCV 2023poster

Standard unsupervised domain adaptation methods adapt models from a source to a target domain using labeled source data and unlabeled target data jointly. In model adaptation, on the other hand, access to the labeled source data is prohibited, i.e., only the source-trained model and unlabeled target…

Cited by 16PDFcodeScholar
2023

Event-Based Frame Interpolation With Ad-Hoc Deblurring

CVPR 2023poster

The performance of video frame interpolation is inherently correlated with the ability to handle motion in the input scene. Even though previous works recognize the utility of asynchronous event information for this task, they ignore the fact that motion may or may not result in blur in the input vi…

2023

Indiscernible Object Counting in Underwater Scenes

CVPR 2023poster

Recently, indiscernible scene understanding has attracted a lot of attention in the vision community. We further advance the frontier of this field by systematically studying a new challenge named indiscernible object counting (IOC), the goal of which is to count objects that are blended with respec…

2023

L2E: Lasers to Events for 6-DoF Extrinsic Calibration of Lidars and Event Cameras

ICRA 2023poster

As neuromorphic technology is maturing, its application to robotics and autonomous vehicle systems has become an area of active research. In particular, event cameras have emerged as a compelling alternative to frame-based cameras in low-power and latency-demanding applications. To enable event came…

Cited by 13SourcecodeScholar
2023

Real-Time Motion Prediction via Heterogeneous Polyline Transformer with Relative Pose Encoding

NeurIPS 2023poster

The real-world deployment of an autonomous driving system requires its components to run on-board and in real-time, including the motion prediction module that predicts the future trajectories of surrounding traffic participants. Existing agent-centric methods have demonstrated outstanding performan…

2023

iDisc: Internal Discretization for Monocular Depth Estimation

CVPR 2023poster

Monocular depth estimation is fundamental for 3D scene understanding and downstream applications. However, even under the supervised setup, it is still challenging and ill-posed due to the lack of geometric constraints. We observe that although a scene can consist of millions of pixels, there are mu…

2022

Event-Based Fusion for Motion Deblurring with Cross-Modal Attention

ECCV 2022poster

"Traditional frame-based cameras inevitably suffer from motion blur due to long exposure times. As a kind of bio-inspired camera, the event camera records the intensity changes in an asynchronous way with high temporal resolution, providing valid image degradation information within the exposure tim…

2022

LiDAR Snowfall Simulation for Robust 3D Object Detection

CVPR 2022oral

3D object detection is a central task for applications such as autonomous driving, in which the system needs to localize and classify surrounding traffic agents, even in the presence of adverse weather. In this paper, we address the problem of LiDAR-based 3D object detection under snowfall. Due to t…

Cited by 147PDFcodeScholar
2022

Lidar Line Selection with Spatially-Aware Shapley Value for Cost-Efficient Depth Completion

CoRL 2022poster

Lidar is a vital sensor for estimating the depth of a scene. Typical spinning lidars emit pulses arranged in several horizontal lines and the monetary cost of the sensor increases with the number of these lines. In this work, we present the new problem of optimizing the positioning of lidar lines to…

Cited by 2SourceScholar
2022

P3Depth: Monocular Depth Estimation With a Piecewise Planarity Prior

CVPR 2022poster

Monocular depth estimation is vital for scene understanding and downstream tasks. We focus on the supervised setup, in which ground-truth depth is available only at training time. Based on knowledge about the high regularity of real 3D scenes, we propose a method that learns to selectively leverage…

Cited by 170PDFcodeScholar
2021

ACDC: The Adverse Conditions Dataset With Correspondences for Semantic Driving Scene Understanding

ICCV 2021poster

Level 5 autonomy for self-driving cars requires a robust visual perception system that can parse input images under any visual condition. However, existing semantic segmentation datasets are either dominated by images captured under normal conditions or are small in scale. To address this, we introd…

Cited by 705PDFScholar
2021

Fog Simulation on Real LiDAR Point Clouds for 3D Object Detection in Adverse Weather

ICCV 2021poster

This work addresses the challenging task of LiDAR-based 3D object detection in foggy weather. Collecting and annotating data in such a scenario is very time, labor and cost intensive. In this paper, we tackle this problem by simulating physically accurate fog into clear-weather scenes, so that the a…

Cited by 185PDFcodeScholar
2019

Guided Curriculum Model Adaptation and Uncertainty-Aware Evaluation for Semantic Nighttime Image Segmentation

ICCV 2019poster

Most progress in semantic segmentation reports on daytime images taken under favorable illumination conditions. We instead address the problem of semantic segmentation of nighttime images and improve the state-of-the-art, by adapting daytime models to nighttime without using nighttime annotations. M…

Cited by 311PDFcodeScholar
2018

Domain Adaptive Faster R-CNN for Object Detection in the Wild

CVPR 2018poster

Object detection typically assumes that training and test data are drawn from an identical distribution, which, however, does not always hold in practice. Such a distribution mismatch will lead to a significant performance drop. In this work, we aim to improve the cross-domain robustness of object d…

2018

Model Adaptation with Synthetic and Real Data for Semantic Dense Foggy Scene Understanding

ECCV 2018poster

This work addresses the problem of semantic scene understanding under dense fog. Although considerable progress has been made in semantic scene understanding, it is mainly related to clear-weather scenes. Extending recognition methods to adverse weather conditions such as fog is crucial for outdoor…

Cited by 285SourcePDFScholar