← Search

Didier Stricker

38 accepted papers

2026

DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance

CVPR 2026

Predicting driver attention is a critical problem for developing explainable autonomous driving systems and understanding driver behavior in mixed human-autonomous vehicle traffic scenarios. Although significant progress has been made through large-scale driver attention datasets and deep learning a

Cited by 0SourcecodeScholar
2026

LiREC-Net: A Target-Free and Learning-Based Network for LiDAR, RGB, and Event Calibration

CVPR 2026

Advanced autonomous systems rely on multi-sensor fusion for safer and more robust perception. To enable effective fusion, calibrating directly from natural driving scenes (i.e., target-free) with high accuracy is crucial for precise multi-sensor alignment. Existing learning-based calibration methods

Cited by 0SourceScholar
2026

NURBGen: High-Fidelity Text-to-CAD Generation Through LLM-Driven NURBS Modeling

AAAI 2026technical

Generating editable 3D CAD models from natural language remains challenging, as existing text-to-CAD systems either produce meshes or rely on scarce design-history data. We present NURBGen, the first framework to generate high fidelity 3D CAD models directly from text using Non-Uniform Rational B-Sp

Cited by 0SourcePDFScholar
2026

ReLaGS: Relational Language Gaussian Splatting

CVPR 2026

Achieving unified 3D perception and reasoning across tasks such as segmentation, retrieval, and relation understanding remains challenging, as existing methods are either object-centric or rely on costly training for inter-object reasoning. We present a novel framework that constructs a hierarchical

Cited by 0SourcecodeScholar
2026

SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking

CVPR 2026

Modeling spinal motion is fundamental to understanding human biomechanics, yet remains underexplored in computer vision due to the spine's complex multi-joint kinematics and the lack of large-scale 3D annotations. We present a biomechanics-aware keypoint simulation framework that augments existing h

Cited by 0SourceScholar
2025

JENGA: Object selection and pose estimation for robotic grasping from a stack

IROS 2025

Vision-based robotic object grasping is typically investigated in the context of isolated objects or unstructured object sets in bin picking scenarios. However, there are several settings, such as construction or warehouse automation, where a robot needs to interact with a structured object formatio

Cited by 0SourceScholar
2025

MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation

CVPR 2025poster

Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an extensive dataset with 40 million text annotations for over 8.9 mi…

2025

STEP-DETR: Advancing DETR-based Semi-Supervised Object Detection with Super Teacher and Pseudo-Label Guided Text Queries

ICCV 2025poster

This paper addresses key limitations in current Semi-Supervised Object Detection (SSOD) frameworks, focusing on issues related to pseudo-label quality, confidence bias, and inefficient query generation. Traditional methods, including CNN-based and DETR-based architectures, often face challenges such…

Cited by 0SourcePDFScholar
2025

TorchAdapt: Towards Light-Agnostic Real-Time Visual Perception

ICCV 2025poster

Low-light conditions significantly degrade the performance of high-level vision tasks. Existing approaches either enhance low-light images without considering normal illumination scenarios, leading to poor generalization, or are tailored to specific tasks. We propose TorchAdapt, a realtime adaptive…

Cited by 0SourcePDFScholar
2024

CLEO: Continual Learning of Evolving Ontologies

ECCV 2024poster

"Continual learning (CL) addresses the problem of catastrophic forgetting in neural networks, which occurs when a trained model tends to overwrite previously learned information, when presented with a new task. CL aims to instill the lifelong learning characteristic of humans in intelligent systems,…

Cited by 0SourcePDFScholar
2024

HiPose: Hierarchical Binary Surface Encoding and Correspondence Pruning for RGB-D 6DoF Object Pose Estimation

CVPR 2024poster

In this work we present a novel dense-correspondence method for 6DoF object pose estimation from a single RGB-D image. While many existing data-driven methods achieve impressive performance they tend to be time-consuming due to their reliance on rendering-based refinement approaches. To circumvent t…

2024

MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding

CVPR 2024poster

3D visual grounding involves matching natural language descriptions with their corresponding objects in 3D spaces. Existing methods often face challenges with accuracy in object recognition and struggle in interpreting complex linguistic queries particularly with descriptions that involve multiple a…

2024

SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and Its Downstream Tasks

CVPR 2024poster

Scene graphs have been recently introduced into 3D spatial understanding as a comprehensive representation of the scene. The alignment between 3D scene graphs is the first step of many downstream tasks such as scene graph aided point cloud registration mosaicking overlap checking and robot navigatio…

2024

Sparse Semi-DETR: Sparse Learnable Queries for Semi-Supervised Object Detection

CVPR 2024poster

In this paper we address the limitations of the DETR-based semi-supervised object detection (SSOD) framework particularly focusing on the challenges posed by the quality of object queries. In DETR-based SSOD the one-to-one assignment strategy provides inaccurate pseudo-labels while the one-to-many a…

Cited by 24SourcePDFScholar
2024

Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text Prompts

NeurIPS 2024spotlight

Prototyping complex computer-aided design (CAD) models in modern softwares can be very time-consuming. This is due to the lack of intelligent systems that can quickly generate simpler intermediate parts. We propose Text2CAD, the first AI framework for generating text-to-parametric CAD models using d…

2023

FeatEnHancer: Enhancing Hierarchical Features for Object Detection and Beyond Under Low-Light Vision

ICCV 2023poster

Extracting useful visual cues for the downstream tasks is especially challenging under low-light vision. Prior works create enhanced representations by either correlating visual quality with machine perception or designing illumination-degrading transformation methods that require pre-training on…

Cited by 35PDFcodeScholar
2023

I2MVFormer: Large Language Model Generated Multi-View Document Supervision for Zero-Shot Image Classification

CVPR 2023highlight

Recent works have shown that unstructured text (documents) from online sources can serve as useful auxiliary information for zero-shot image classification. However, these methods require access to a high-quality source like Wikipedia and are limited to a single source of information. Large Language…

2023

Introducing Language Guidance in Prompt-based Continual Learning

ICCV 2023poster

Continual Learning aims to learn a single model on a sequence of tasks without having access to data from previous tasks. The biggest challenge in the domain still remains catastrophic forgetting: a loss in performance on seen classes of earlier tasks. Some existing methods rely on an expensive repl…

Cited by 56PDFcodeScholar
2023

OPA-3D: Occlusion-Aware Pixel-Wise Aggregation for Monocular 3D Object Detection

RA-L 2023

Monocular 3D object detection has recently made a significant leap forward thanks to the use of pre-trained depth estimators for pseudo-LiDAR recovery. Yet, such two-stage methods typically suffer from overfitting and are incapable of explicitly encapsulating the geometric relation between depth and

Cited by 39SourceScholar
2023

Structure PLP-SLAM: Efficient Sparse Mapping and Localization using Point, Line and Plane for Monocular, RGB-D and Stereo Cameras

ICRA 2023poster

This paper presents a visual SLAM system that uses both points and lines for robust camera localization, and simultaneously performs a piece-wise planar reconstruction (PPR) of the environment to provide a structural map in real-time. One of the biggest challenges in parallel tracking and mapping wi…

Cited by 70SourcecodeScholar
2023

U-RED: Unsupervised 3D Shape Retrieval and Deformation for Partial Point Clouds

ICCV 2023poster

In this paper, we propose U-RED, an Unsupervised shape REtrieval and Deformation pipeline that takes an arbitrary object observation as input, typically captured by RGB images or scans, and jointly retrieves and deforms the geometrically similar CAD models from a pre-established database to tightly…

Cited by 3PDFcodeScholar
2022

RMS-FlowNet: Efficient and Robust Multi-Scale Scene Flow Estimation for Large-Scale Point Clouds

ICRA 2022poster

The proposed RMS-FlowNet is a novel end-to-end learning-based architecture for accurate and efficient scene flow estimation which can operate on point clouds of high density. For hierarchical scene flow estimation, the existing methods depend on either expensive Farthest-Point-Sampling (FPS) or stru…

Cited by 22SourceScholar
2022

SOMSI: Spherical Novel View Synthesis With Soft Occlusion Multi-Sphere Images

CVPR 2022poster

Spherical novel view synthesis (SNVS) is the task of estimating 360 views at dynamic novel views given a set of 360 input views. Prior arts learn multi-sphere image (MSI) representations that enables fast rendering times but are only limited to modelling low-dimensional color values. Modelling high-…

Cited by 17PDFcodeScholar
2022

Towards Artefact Aware Human Motion Capture using Inertial Sensors Integrated into Loose Clothing

ICRA 2022poster

Inertial motion capture has become an attractive alternative to optical motion capture for human joint angle estimation outside the laboratory. Usually inertial sensors are assumed to be tightly fixed to the body segments, which can be cumbersome regarding setup-time and ease-of-use. However, integr…

Cited by 16SourceScholar
2022

ZebraPose: Coarse To Fine Surface Encoding for 6DoF Object Pose Estimation

CVPR 2022poster

Establishing correspondences from image to 3D has been a key task of 6DoF object pose estimation for a long time. To predict pose more accurately, deeply learned dense maps replaced sparse templates. Dense methods also improved pose estimation in the presence of occlusion. More recently researchers…

Cited by 176PDFcodeScholar
2021

PlaneSegNet: Fast and Robust Plane Estimation Using a Single-stage Instance Segmentation CNN

ICRA 2021poster

Instance segmentation of planar regions in indoor scenes benefits visual SLAM and other applications such as augmented reality (AR) where scene understanding is required. Existing methods built upon two-stage frameworks show satisfactory accuracy but are limited by low frame rates. In this work, we…

Cited by 16SourceScholar
2021

RPSRNet: End-to-End Trainable Rigid Point Set Registration Network Using Barnes-Hut 2D-Tree Representation

CVPR 2021poster

We propose RPSRNet - a novel end-to-end trainable deep neural network for rigid point set registration. For this task, we use a novel 2^D-tree representation for the input point sets and a hierarchical deep feature embedding in the neural network. An iterative transformation refinement module in our…

Cited by 20PDFScholar
2020

DeepLiDARFlow: A Deep Learning Architecture For Scene Flow Estimation Using Monocular Camera and Sparse LiDAR

IROS 2020poster

Scene flow is the dense 3D reconstruction of motion and geometry of a scene. Most state-of-the-art methods use a pair of stereo images as input for full scene reconstruction. These methods depend a lot on the quality of the RGB images and perform poorly in regions with reflective objects, shadows, i…

Cited by 38SourcecodeScholar
2020

Generative View Synthesis: From Single-view Semantics to Novel-view Images

NeurIPS 2020poster

Content creation, central to applications such as virtual reality, can be tedious and time-consuming. Recent image synthesis methods simplify this task by offering tools to generate new views from as little as a single input image, or by converting a semantic map into a photorealistic image. We pro…

2020

HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth Map

CVPR 2020poster

3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. The state-of-the-art methods directly regress 3D hand meshes from 2D depth images via 2D convolutional neural networks, which leads to artefacts in the estimations due t…

Cited by 93PDFScholar
2019

Accelerated Gravitational Point Set Alignment With Altered Physical Laws

ICCV 2019poster

This work describes Barnes-Hut Rigid Gravitational Approach (BH-RGA) -- a new rigid point set registration method relying on principles of particle dynamics. Interpreting the inputs as two interacting particle swarms, we directly minimise the gravitational potential energy of the system using non-li…

Cited by 16PDFcodeScholar
2019

LiDAR-Flow: Dense Scene Flow Estimation from Sparse LiDAR and Stereo Images

IROS 2019poster

We propose a new approach called LiDAR-Flow to robustly estimate a dense scene flow by fusing a sparse LiDAR with stereo images. We take the advantage of the high accuracy of LiDAR to resolve the lack of information in some regions of stereo images due to textureless objects, shadows, ill-conditione…

Cited by 35SourceScholar
2019

SDC - Stacked Dilated Convolution: A Unified Descriptor Network for Dense Matching Tasks

CVPR 2019oral

Dense pixel matching is important for many computer vision tasks such as disparity and flow estimation. We present a robust, unified descriptor network that considers a large context region with high spatial variance. Our network has a very large receptive field and avoids striding layers to mainta…

Cited by 53PDFcodeScholar
2018

Learning 3D Shapes as Multi-Layered Height-maps using 2D Convolutional Networks

ECCV 2018poster

We present a novel global representation of 3D shapes, suitable for the application of 2D CNNs. We represent 3D shapes as multi-layered height maps (MLH) where at each grid location, we store multiple instances of height maps, thereby representing 3D shape detail that is hidden behind several layers…

Cited by 40SourcePDFScholar
2017

CNN-Based Patch Matching for Optical Flow With Thresholded Hinge Embedding Loss

CVPR 2017poster

Learning based approaches have not yet achieved their full potential in optical flow estimation, where their performance still trails heuristic approaches. In this paper, we present a CNN based patch matching approach for optical flow estimation. An important contribution of our approach is a novel…

Cited by 149PDFScholar
2015

Flow Fields: Dense Correspondence Fields for Highly Accurate Large Displacement Optical Flow Estimation

ICCV 2015oral

Modern large displacement optical flow algorithms usually use an initialization by either sparse descriptor matching techniques or dense approximate nearest neighbor fields. While the latter have the advantage of being dense, they have the major disadvantage of being very outlier prone as they are n…

Cited by 325PDFScholar