← Search

Markus Gross

35 accepted papers

2026

CoCo-InEKF: State Estimation with Learned Contact Covariances in Dynamic, Contact-Rich Scenarios

RSS 2026poster

Robust state estimation for highly dynamic motion of legged robots remains challenging, especially in dynamic, contact-rich scenarios. Traditional approaches often rely on binary contact states that fail to capture the nuances of partial contact or directional slippage. This paper presents CoCo-InEK…

Cited by 0SourceScholar
2026

Efficient All-Pairs Correlation Volume Sampling for Optical Flow Estimation

CVPR 2026

Recent optical flow estimation methods often employ local cost sampling from a dense all-pairs correlation volume. This results in quadratic computational and memory complexity in the number of pixels. Although an alternative memory-efficient implementation with on-demand cost computation exists, th

Cited by 0SourceScholar
2026

EnerGS: Energy-Based Gaussian Splatting under Partial Geometric Observability

ICML 2026poster

3D Gaussian Splatting (3DGS) has been widely adopted for scene reconstruction, where training inherently constitutes a highly coupled and non-convex optimization problem. Recent works commonly incorporate geometric priors, such as LiDAR measurements, either for initialization or as training constrai…

Cited by 0SourceScholar
2026

Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel Views

CVPR 2026

Soft boundaries, like thin hairs, are commonly observed in natural and computer-generated imagery, but they remain challenging for 3D vision due to the ambiguous mixing of foreground and background cues. This paper introduces Guardians of the Hair (HairGuard), a framework designed to recover fine-gr

Cited by 0SourceScholar
2026

OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective

CVPR 2026

Semantic Scene Completion (SSC) is essential for 3D perception in mobile robotics, as it enables holistic scene understanding by jointly estimating dense volumetric occupancy and per-voxel semantics. Although SSC has been widely studied in terrestrial domains such as autonomous driving, aerial setti

Cited by 0SourcecodeScholar
2026

RelightAnyone: A Generalized Relightable 3D Gaussian Head Model

CVPR 2026

3D Gaussian Splatting (3DGS) has become a standard approach to reconstruct and render photorealistic 3D head avatars. A major challenge is to relight the avatars to match any scene illumination. For high quality relighting, existing methods require subjects to be captured under complex time-multiple

Cited by 0SourceScholar
2026

What Is It Like to Be a Noise? An Entropy-based Gaussian Noise Regularization for Diffusion Models

CVPR 2026

Inference-time optimization of diffusion latents enables powerful control but often degrades the statistical structure of true Gaussian noise, causing artifacts and reward hacking. To address this, we propose a Gaussianity regularizer that aligns a sample's local statistics with a typical Gaussian r

Cited by 0SourceScholar
2025

Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image Compression

CVPR 2025poster

Generative neural image compression supports data representation with extremely low bitrate, allowing clients to synthesize details and consistently producing highly realistic images. By leveraging the similarities between quantization error and additive noise, diffusion-based generative image compr…

Cited by 0SourcePDFScholar
2025

IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals

NeurIPS 2025poster

Semantic Scene Completion (SSC) has emerged as a pivotal approach for jointly learning scene geometry and semantics, enabling downstream applications such as navigation in mobile robotics. The recent generalization to Panoptic Scene Completion (PSC) advances the SSC domain by integrating instance-le…

Cited by 0SourceScholar
2025

LDIP: Long Distance Information Propagation for Video Super-Resolution

ICCV 2025poster

Video super-resolution (VSR) methods typically exploit information across multiple frames to achieve high quality upscaling, with recent approaches demonstrating impressive performance. Nevertheless, challenges remain, particularly in effectively leveraging information over long distances. To addres…

Cited by 0SourcePDFScholar
2025

LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping

CVPR 2025poster

Anamorphosis refers to a category of images that are intentionally distorted, making them unrecognizable when viewed directly. Their true form only reveals itself when seen from a specific viewpoint, which can be through some catadioptric device like a mirror or a lens. While the construction of the…

Cited by 0SourcePDFScholar
2025

Monocular Facial Appearance Capture in the Wild

ICCV 2025poster

We present a new method for reconstructing the appearance properties of human faces from a lightweight capture procedure in an unconstrained environment. Our method recovers the surface geometry, diffuse albedo, specular intensity and specular roughness from a monocular video containing a simple hea…

Cited by 0SourcePDFScholar
2024

Artist-Friendly Relightable and Animatable Neural Heads

CVPR 2024poster

An increasingly common approach for creating photo-realistic digital avatars is through the use of volumetric neural fields. The original neural radiance field (NeRF) allowed for impressive novel view synthesis of static heads when trained on a set of multi-view images and follow up methods showed t…

Cited by 4SourcePDFScholar
2024

BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation

NeurIPS 2024poster

By training over large-scale datasets, zero-shot monocular depth estimation (MDE) methods show robust performance in the wild but often suffer from insufficient detail. Although recent diffusion-based MDE approaches exhibit a superior ability to extract details, they struggle in geometrically comple…

Cited by 7SourcePDFScholar
2024

Collaborative Semantic Occupancy Prediction with Hybrid Feature Fusion in Connected Automated Vehicles

CVPR 2024poster

Collaborative perception in automated vehicles leverages the exchange of information between agents aiming to elevate perception results. Previous camera-based collaborative 3D perception methods typically employ 3D bounding boxes or bird's eye views as representations of the environment. However th…

Cited by 19SourcePDFScholar
2024

How I Warped Your Noise: a Temporally-Correlated Noise Prior for Diffusion Models

ICLR 2024oral

Video editing and generation methods often rely on pre-trained image-based diffusion models. During the diffusion process, however, the reliance on rudimentary noise sampling techniques that do not preserve correlations present in subsequent frames of a video is detrimental to the quality of the res…

Cited by 16SourcePDFScholar
2024

Lossy Image Compression with Foundation Diffusion Models

ECCV 2024poster

"Incorporating diffusion models in the image compression domain has the potential to produce realistic and detailed reconstructions, especially at extremely low bitrates. Previous methods focus on using diffusion models as expressive decoders robust to quantization errors in the conditioning signals…

Cited by 11SourcePDFScholar
2024

QUADify: Extracting Meshes with Pixel-level Details and Materials from Images

CVPR 2024highlight

Despite exciting progress in automatic 3D reconstruction from images excessive and irregular triangular faces in the resulting meshes still constitute a significant challenge when it comes to adoption in practical artist workflows. Therefore we propose a method to extract regular quad-dominant meshe…

Cited by 0SourcePDFScholar
2023

Frame Interpolation Transformer and Uncertainty Guidance

CVPR 2023poster

Video frame interpolation has seen important progress in recent years, thanks to developments in several directions. Some works leverage better optical flow methods with improved splatting strategies or additional cues from depth, while others have investigated alternative approaches through direct…

Cited by 15SourcePDFScholar
2023

Kernel Aware Resampler

CVPR 2023poster

Deep learning based methods for super-resolution have become state-of-the-art and outperform traditional approaches by a significant margin. From the initial models designed for fixed integer scaling factors (e.g. x2 or x4), efforts were made to explore different directions such as modeling blur ker…

Cited by 3SourcePDFScholar
2023

ReNeRF: Relightable Neural Radiance Fields with Nearfield Lighting

ICCV 2023poster

Recent work on radiance fields and volumetric inverse rendering (e.g., NeRFs) has provided excellent results in building data-driven models of real scenes for novel view synthesis with high photorealism. While full control over viewpoint is achieved, scene lighting is typically "baked" into the mode…

Cited by 21PDFScholar
2023

Transformer-Based Neural Augmentation of Robot Simulation Representations

RA-L 2023

Simulation representations of robots have advanced in recent years. Yet, there remain significant sim-to-real gaps because of modeling assumptions and hard-to-model behaviors such as friction. In this letter, we propose to augment common simulation representations with a transformer-inspired archite

Cited by 6SourceScholar
2021

Adaptive Convolutions for Structure-Aware Style Transfer

CVPR 2021poster

Style transfer between images is an artistic application of CNNs, where the 'style' of one image is transferred onto another image while preserving the latter's content. The state of the art in neural style transfer is based on Adaptive Instance Normalization (AdaIN), a technique that transfers the…

Cited by 81PDFScholar
2020

Attention-Driven Cropping for Very High Resolution Facial Landmark Detection

CVPR 2020poster

Facial landmark detection is a fundamental task for many consumer and high-end applications and is almost entirely solved by machine learning methods today. Existing datasets used to train such algorithms are primarily made up of only low resolution images, and current algorithms are limited to inpu…

Cited by 87PDFScholar
2019

Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Value Approximation

ICML 2019oral

The problem of explaining the behavior of deep neural networks has recently gained a lot of attention. While several attribution methods have been proposed, most come without strong theoretical foundations, which raises questions about their reliability. On the other hand, the literature on cooperat…

Cited by 319SourcePDFScholar
2019

Learning-Based Sampling for Natural Image Matting

CVPR 2019poster

The goal of natural image matting is the estimation of opacities of a user-defined foreground object that is essential in creating realistic composite imagery. Natural matting is a challenging process due to the high number of unknowns in the mathematical modeling of the problem, namely the opacitie…

Cited by 160PDFScholar
2018

A Network Architecture for Point Cloud Classification via Automatic Depth Images Generation

CVPR 2018poster

We propose a novel neural network architecture for point cloud classification. Our key idea is to automatically transform the 3D unordered input data into a set of useful 2D depth images, and classify them by exploiting well performing image classification CNNs. We present new differentiable module…

Cited by 82SourcePDFScholar
2018

A Neural Multi-Sequence Alignment TeCHnique (NeuMATCH)

CVPR 2018poster

The alignment of heterogeneous sequential data (video to text) is an important and challenging problem. Standard techniques for this task, including Dynamic Time Warping (DTW) and Conditional Random Fields (CRFs), suffer from inherent drawbacks. Mainly, the Markov assumption implies that, given the…

2018

PhaseNet for Video Frame Interpolation

CVPR 2018poster

Most approaches for video frame interpolation require accurate dense correspondences to synthesize an in-between frame. Therefore, they do not perform well in challenging scenarios with e.g. lighting changes or motion blur. Recent deep learning approaches that rely on kernels to represent motion can…

Cited by 230SourcePDFScholar
2018

Towards better understanding of gradient-based attribution methods for Deep Neural Networks

ICLR 2018poster

Understanding the flow of information in Deep Neural Networks (DNNs) is a challenging problem that has gain increasing attention over the last few years. While several methods have been proposed to explain network predictions, there have been only a few attempts to compare them from a theoretical pe…

2017

Human Shape From Silhouettes Using Generative HKS Descriptors and Cross-Modal Neural Networks

CVPR 2017spotlight

In this work, we present a novel method for capturing human body shape from a single scaled silhouette. We combine deep correlated features capturing different 2D views, and embedding spaces based on 3D cues in a novel convolutional neural network (CNN) based architecture. We first train a CNN to fi…

Cited by 127PDFScholar
2016

A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation

CVPR 2016poster

Over the years, datasets and benchmarks have proven their fundamental importance in computer vision research, enabling targeted progress and objective comparisons in many fields. At the same time, legacy datasets may impend the evolution of a field due to saturated algorithm performance and the lack…

Cited by 2422PDFcodeScholar
2015

Fully Connected Object Proposals for Video Segmentation

ICCV 2015poster

We present a novel approach to video segmentation using multiple object proposals. The problem is formulated as a minimization of a novel energy function defined over a fully connected graph of object proposals. Our model combines appearance with long-range point tracks, which is key to ensure robus…

Cited by 210PDFScholar