← Search

Gerhard Rigoll

6 accepted papers

2026

SpaRC: Sparse Radar-Camera Fusion for 3D Object Detection

ICRA 2026poster

In this work, we present SpaRC, a novel sparse fusion transformer for 3D perception that integrates multi-view image semantics with Radar and Camera point features. The fusion of radar and camera modalities has emerged as an efficient perception paradigm for autonomous driving systems. While convent…

2025

Optimizing Robot Programming: Mixed Reality Gripper Control

ICRA 2025

Conventional robot programming methods are complex and time-consuming for users. In recent years, alternative approaches such as mixed reality have been explored to address these challenges and optimize robot programming. While the findings of the mixed reality robot programming methods are convinci

Cited by 0SourceScholar
2025

Unleashing HyDRa: Hybrid Fusion, Depth Consistency and Radar for Unified 3D Perception

ICRA 2025

Low-cost, vision-centric 3D perception systems for autonomous driving have made significant progress in recent years, narrowing the gap to expensive LiDAR-based methods. The primary challenge in becoming a fully reliable alternative lies in robust depth prediction capabilities, as camera-based syste

Cited by 43SourcecodeScholar
2021

How To Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild

ICCV 2021poster

Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and (iii) temporal modeling for the reference speaker. Each sta…

Cited by 62PDFcodeScholar
2020

Small-Footprint Keyword Spotting on Raw Audio Data with Sinc-Convolutions

ICASSP 2020accepted

Keyword Spotting (KWS) enables speech-based user interaction on smart devices. Always-on and battery-powered application scenarios for smart devices put constraints on hardware resources and power consumption, while also demanding high accuracy as well as real-time capability. Previous architectures…

Cited by 0SourceScholar
2018

Robust Facial Landmark Detection via a Fully-Convolutional Local-Global Context Network

CVPR 2018poster

While fully-convolutional neural networks are very strong at modeling local features, they fail to aggregate global context due to their constrained receptive field. Modern methods typically address the lack of global context by introducing cascades, pooling, or by fitting a statistical model. In th…

Cited by 112SourcePDFScholar