← Search

Vladislav Golyanik

52 accepted papers

2026

$\boldsymbol{\partial^\infty}$-Grid: A Neural Differential Equation Solver with Differentiable Feature Grids

ICLR 2026poster

We present a novel differentiable grid-based representation for efficiently solving differential equations (DEs). Widely used architectures for neural solvers, such as sinusoidal neural networks, are coordinate-based MLPs that are, both, computationally intensive and slow to train. Although grid-bas…

Cited by 0SourcecodeScholar
2026

E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimation

CVPR 2026

Event cameras offer multiple advantages in monocular egocentric 3D human pose estimation from head-mounted devices, such as millisecond temporal resolution, high dynamic range, and negligible motion blur. Existing methods effectively leverage these properties, but suffer from low 3D estimation accur

Cited by 0SourceScholar
2026

Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views

CVPR 2026

We present _Relightable Holoported Characters_ (RHC), a novel person-specific method for free-view rendering and relighting of full-body and highly dynamic humans solely observed from sparse-view RGB videos at inference. In contrast to classical one-light-at-a-time (OLAT)-based human relighting, our

Cited by 0SourceScholar
2026

SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokens

CVPR 2026

Synthesizing text-driven 3D human motion within realistic scenes requires learning both semantic intent ("walk to the couch") and physical feasibility (e.g., avoiding collisions). Current methods use generative frameworks that simultaneously learn high-level planning and low-level contact reasoning,

Cited by 0SourcecodeScholar
2026

Splat the Net: Radiance Fields with Splattable Neural Primitives

ICLR 2026poster

Radiance fields have emerged as a predominant representation for modeling 3D scene appearance. Neural formulations such as Neural Radiance Fields provide high expressivity but require costly ray marching for rendering, whereas primitive-based methods such as 3D Gaussian Splatting offer real-time eff…

Cited by 0SourcecodeScholar
2025

Attention (as Discrete-Time Markov) Chains

NeurIPS 2025poster

We introduce a new interpretation of the attention matrix as a discrete-time Markov chain. Our interpretation sheds light on common operations involving attention scores such as selection, summation, and averaging in a unified framework. It further extends them by considering indirect attention, pro…

Cited by 0SourceScholar
2025

BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects

CVPR 2025poster

We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-base…

Cited by 2SourcePDFScholar
2025

Bring Your Rear Cameras for Egocentric 3D Human Pose Estimation

ICCV 2025poster

Egocentric 3D human pose estimation has been actively studied using cameras installed in front of a head-mounted device (HMD). While frontal placement is the optimal and the only option for some tasks, such as hand tracking, it remains unclear if the same holds for full-body tracking due to self-occ…

2025

DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image

ICLR 2025poster

Reconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, co…

2025

Denoising Functional Maps: Diffusion Models for Shape Correspondence

CVPR 2025poster

Estimating correspondences between pairs of deformable shapes remains a challenging problem. Despite substantial progress, existing methods lack broad generalization capabilities and require category-specific training data. To address these limitations, we propose a fundamentally new approach to sha…

2025

HumanOLAT: A Large-Scale Dataset for Full-Body Human Relighting and Novel-View Synthesis

ICCV 2025poster

Simultaneous relighting and novel-view rendering of digital human representations is an important yet challenging task with numerous applications. However, progress in this area has been significantly limited due to the lack of publicly available, high-quality datasets, especially for full-body huma…

Cited by 0SourcePDFScholar
2025

QuCOOP: A Versatile Framework for Solving Composite and Binary-Parametrised Problems on Quantum Annealers

CVPR 2025highlight

There is growing interest in solving computer vision problems such as mesh or point set alignment using Adiabatic Quantum Computing (AQC). Unfortunately, modern experimental AQC devices such as D-Wave only support Quadratic Unconstrained Binary Optimisation (QUBO) problems, which severely limits the…

2025

Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation Fields

CVPR 2025poster

3D reconstruction of highly deformable surfaces (e.g. cloths) from monocular RGB videos is a challenging problem, and no solution provides a consistent and accurate recovery of fine-grained surface details. To account for the ill-posed nature of the setting, existing methods use deformation models w…

Cited by 0SourcePDFScholar
2024

3D Human Pose Perception from Egocentric Stereo Videos

CVPR 2024highlight

While head-mounted devices are becoming more compact they provide egocentric views with significant self-occlusions of the device user. Hence existing methods often fail to accurately estimate complex 3D poses from egocentric views. In this work we propose a new transformer-based framework to improv…

Cited by 19SourcePDFScholar
2024

EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams

CVPR 2024poster

Monocular egocentric 3D human motion capture is a challenging and actively researched problem. Existing methods use synchronously operating visual sensors (e.g. RGB cameras) and often fail under low lighting and fast motions which can be restricting in many applications involving head-mounted device…

2024

Holoported Characters: Real-time Free-viewpoint Rendering of Humans from Sparse RGB Cameras

CVPR 2024poster

We present the first approach to render highly realistic free-viewpoint videos of a human actor in general apparel from sparse multi-view recording to display in real-time at an unprecedented 4K resolution. At inference our method only requires four camera views of the moving actor and the respectiv…

Cited by 9SourcePDFScholar
2024

NeuralClothSim: Neural Deformation Fields Meet the Thin Shell Theory

NeurIPS 2024poster

Despite existing 3D cloth simulators producing realistic results, they predominantly operate on discrete surface representations (e.g. points and meshes) with a fixed spatial resolution, which often leads to large memory consumption and resolution-dependent simulations. Moreover, back-propagating gr…

2024

ReMoS: 3D Motion-Conditioned Reaction Synthesis for Two-Person Interactions

ECCV 2024poster

"Current approaches for 3D human motion synthesis generate high-quality animations of digital humans performing a wide variety of actions and gestures. However, a notable technological gap exists in addressing the complex dynamics of multi-human interactions within this paradigm. In this work, we pr…

2024

Relightable Neural Actor with Intrinsic Decomposition and Pose Control

ECCV 2024poster

"Creating a controllable and relightable digital avatar from multi-view video with fixed illumination is a very challenging problem since humans are highly articulated, creating pose-dependent appearance effects, and skin as well as clothing require space-varying BRDF modeling. Existing works on cre…

Cited by 4SourcePDFScholar
2024

VINECS: Video-based Neural Character Skinning

CVPR 2024poster

Rigging and skinning clothed human avatars is a challenging task and traditionally requires a lot of manual work and expertise. Recent methods addressing it either generalize across different characters or focus on capturing the dynamics of a single character observed under different pose configurat…

Cited by 3SourcePDFScholar
2023

CCuantuMM: Cycle-Consistent Quantum-Hybrid Matching of Multiple Shapes

CVPR 2023poster

Jointly matching multiple, non-rigidly deformed 3D shapes is a challenging, NP-hard problem. A perfect matching is necessarily cycle-consistent: Following the pairwise point correspondences along several shapes must end up at the starting vertex of the original shape. Unfortunately, existing quantum…

Cited by 15SourcePDFScholar
2023

EventNeRF: Neural Radiance Fields From a Single Colour Event Camera

CVPR 2023poster

Asynchronously operating event cameras find many applications due to their high dynamic range, vanishingly low motion blur, low latency and low data bandwidth. The field saw remarkable progress during the last few years, and existing event-based 3D reconstruction approaches recover sparse point clou…

2023

Mofusion: A Framework for Denoising-Diffusion-Based Motion Synthesis

CVPR 2023highlight

Conventional methods for human motion synthesis have either been deterministic or have had to struggle with the trade-off between motion diversity vs motion quality. In response to these limitations, we introduce MoFusion, i.e., a new denoising-diffusion-based framework for high-quality conditional…

Cited by 187SourcePDFScholar
2023

QuAnt: Quantum Annealing with Learnt Couplings

ICLR 2023top-25%

Modern quantum annealers can find high-quality solutions to combinatorial optimisation objectives given as quadratic unconstrained binary optimisation (QUBO) problems. Unfortunately, obtaining suitable QUBO forms in computer vision remains challenging and currently requires problem-specific analytic…

Cited by 5SourcePDFScholar
2023

Quantum Multi-Model Fitting

CVPR 2023highlight

Geometric model fitting is a challenging but fundamental computer vision problem. Recently, quantum optimization has been shown to enhance robust fitting for the case of a single model, while leaving the question of multi-model fitting open. In response to this challenge, this paper shows that the l…

2023

Self-Supervised Pre-Training With Masked Shape Prediction for 3D Scene Understanding

CVPR 2023poster

Masked signal modeling has greatly advanced self-supervised pre-training for language and 2D images. However, it is still not fully explored in 3D scene understanding. Thus, this paper introduces Masked Shape Prediction (MSP), a new framework to conduct masked signal modeling in 3D scenes. MSP uses…

2022

HULC: 3D HUman Motion Capture with Pose Manifold SampLing and Dense Contact Guidance

ECCV 2022poster

"Marker-less monocular 3D human motion capture (MoCap) with scene interactions is a challenging research topic relevant for extended reality, robotics and virtual avatar generation. Due to the inherent depth ambiguity of monocular settings, 3D motions captured with existing methods often contain sev…

Cited by 31SourcePDFScholar
2022

NeRF for Outdoor Scene Relighting

ECCV 2022poster

"Photorealistic editing of outdoor scenes from photographs requires a profound understanding of the image formation process and an accurate estimation of the scene geometry, reflectance and illumination. A delicate manipulation of the lighting can then be performed while keeping the scene albedo and…

Cited by 150SourcePDFScholar
2022

Physical Inertial Poser (PIP): Physics-Aware Real-Time Human Motion Tracking From Sparse Inertial Sensors

CVPR 2022poster

Motion capture from sparse inertial sensors has shown great potential compared to image-based approaches since occlusions do not lead to a reduced tracking quality and the recording space is not restricted to be within the viewing frustum of the camera. However, capturing the motion and global posit…

Cited by 200PDFScholar
2022

Playable Environments: Video Manipulation in Space and Time

CVPR 2022poster

We present Playable Environments - a new representation for interactive video generation and manipulation in space and time. With a single image at inference time, our novel framework allows the user to move objects in 3D while generating a video by providing a sequence of desired actions. The actio…

Cited by 20PDFcodeScholar
2022

Q-FW: A Hybrid Classical-Quantum Frank-Wolfe for Quadratic Binary Optimization

ECCV 2022poster

"We present a hybrid classical-quantum framework based on the Frank-Wolfe algorithm, Q-FW, for solving quadratic, linearly-constrained, binary optimization problems on quantum annealers (QA). The computational premise of quantum computers has cultivated the re-design of various existing vision probl…

2022

Quantum Motion Segmentation

ECCV 2022poster

Motion segmentation is a challenging problem that seeks to identify independent motions in two or several input images. This paper introduces the first algorithm for motion segmentation that relies on adiabatic quantum optimization of the objective function. The proposed method achieves on-par perfo…

2022

UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture

ECCV 2022poster

"We present UnrealEgo, a new large-scale naturalistic dataset for egocentric 3D human pose estimation. UnrealEgo is based on an advanced concept of eyeglasses equipped with two fisheye cameras that can be used in unconstrained environments. We design their virtual prototype and attach them to 3D hum…

Cited by 54SourcePDFScholar
2022

f-SfT: Shape-From-Template With a Physics-Based Deformation Model

CVPR 2022poster

Shape-from-Template (SfT) methods estimate 3D surface deformations from a single monocular RGB camera while assuming a 3D state known in advance (a template). This is an important yet challenging problem due to the under-constrained nature of the monocular setting. Existing SfT techniques predominan…

Cited by 24PDFcodeScholar
2021

EventHands: Real-Time Neural 3D Hand Pose Estimation From an Event Stream

ICCV 2021poster

3D hand pose estimation from monocular videos is a long-standing and challenging problem, which is now seeing a strong upturn. In this work, we address it for the first time using a single event camera, i.e., an asynchronous vision sensor reacting on brightness changes. Our EventHands approach has c…

Cited by 62PDFcodeScholar
2021

Gravity-Aware Monocular 3D Human-Object Reconstruction

ICCV 2021poster

This paper proposes GraviCap, i.e., a new approach for joint markerless 3D human motion capture and object trajectory estimation from monocular RGB videos. We focus on scenes with objects partially observed during a free flight. In contrast to existing monocular methods, we can recover scale, object…

Cited by 31PDFScholar
2021

High-Fidelity Neural Human Motion Transfer From Monocular Video

CVPR 2021poster

Video-based human motion transfer creates video animations of humans following a source motion. Current methods show remarkable results for tightly-clad subjects. However, the lack of temporally consistent handling of plausible clothing dynamics, including fine and high-frequency details, significan…

Cited by 41PDFScholar
2021

Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular Video

ICCV 2021poster

We present Non-Rigid Neural Radiance Fields (NR-NeRF), a reconstruction and novel view synthesis approach for general non-rigid dynamic scenes. Our approach takes RGB images of a dynamic scene as input (e.g., from a monocular video recording), and creates a high-quality space-time geometry and appea…

Cited by 557PDFScholar
2021

Pose-Guided Human Animation From a Single Image in the Wild

CVPR 2021poster

We present a new pose transfer method for synthesizing a human animation from a single image of a person controlled by a sequence of body poses. Existing pose transfer methods exhibit significant visual artifacts when applying to a novel scene, resulting in temporal inconsistency and failures in pre…

Cited by 77PDFScholar
2021

Q-Match: Iterative Shape Matching via Quantum Annealing

ICCV 2021poster

Finding shape correspondences can be formulated as an NP-hard quadratic assignment problem (QAP) that becomes infeasible for shapes with high sampling density. A promising research direction is to tackle such quadratic optimization problems over binary variables with quantum annealing, which allows…

Cited by 37PDFScholar
2020

DEMEA: Deep Mesh Autoencoders for Non-Rigidly Deforming Objects

ECCV 2020poster

Mesh autoencoders are commonly used for dimensionality reduction, sampling and mesh modeling. We propose a general-purpose DEep MEsh Autoencoder \hbox{(DEMEA)} which adds a novel embedded deformation layer to a graph-convolutional mesh autoencoder. The embedded deformation layer (EDL) is a different…

Cited by 51SourcePDFScholar
2020

EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event Camera

CVPR 2020oral

The high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement, the high data bandwidth and the consequent high computation overhead. In this paper, we propose EventCap -- the first…

Cited by 124PDFScholar
2020

HTML: A Parametric Hand Texture Model for 3D Hand Reconstruction and Personalization

ECCV 2020poster

3D hand reconstruction from images is a widely-studied problem in computer vision and graphics, and has a particularly high relevance for virtual and augmented reality. Although several 3D hand reconstruction approaches leverage hand models as a strong prior to resolve ambiguities and achieve more r…

Cited by 87SourcePDFScholar
2020

HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth Map

CVPR 2020poster

3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. The state-of-the-art methods directly regress 3D hand meshes from 2D depth images via 2D convolutional neural networks, which leads to artefacts in the estimations due t…

Cited by 93PDFScholar
2020

Neural Dense Non-Rigid Structure from Motion with Latent Space Constraints

ECCV 2020poster

We introduce the first dense neural non-rigid structure from motion (N-NRSfM) approach, which can be trained end-to-end in an unsupervised manner from 2D point tracks. Compared to the competing methods, our combination of loss functions is fully-differentiable and can be readily integrated into deep…

Cited by 66SourcePDFScholar
2020

Neural Re-Rendering of Humans from a Single Image

ECCV 2020poster

Human re-rendering from a single image is a starkly under-constrained problem and state-of-the-art algorithms often exhibit un-desired artefacts, such as oversmoothing, unrealistic distortions of thebody parts and garments, or implausible changes of the texture. To ad-dress these challenges, we prop…

Cited by 91SourcePDFScholar
2020

PatchNets: Patch-Based Generalizable Deep Implicit 3D Shape Representations

ECCV 2020poster

Implicit surface representation combined with deep learning has led to impressive models which can represent detailed shapes of objects. Implicit surface representations, such as signed-distance functions, allow to represent shapes of arbitrary topologies. Since a continous function is learned, the…

Cited by 116SourcePDFScholar
2019

Accelerated Gravitational Point Set Alignment With Altered Physical Laws

ICCV 2019poster

This work describes Barnes-Hut Rigid Gravitational Approach (BH-RGA) -- a new rigid point set registration method relying on principles of particle dynamics. Interpreting the inputs as two interacting particle swarms, we directly minimise the gravitational potential energy of the system using non-li…

Cited by 16PDFcodeScholar