← Search

Klemen Kotar

11 accepted papers

2026

Perceptual 3D Simulation With Physical World Modeling

CVPR 2026

Predicting how a scene will evolve after a desired 3D transformation from images is a central goal in vision, graphics, and robotics. Yet unlike ideal simulators with full access to 3D geometry and dynamics, real world systems must rely on perceptual inputs and local actions that are inherently part

Cited by 0SourceScholar
2026

Physical Object Understanding with a Physically Controllable World Model

CVPR 2026

A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solving these tasks requires world models capable of inferring distributional states of the world from partial observations -

Cited by 0SourceScholar
2026

Unified 3D Scene Understanding Through Physical World Modeling

ICLR 2026poster

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Existing approaches have typically addressed these tasks in isolation, preventing th…

Cited by 0SourceScholar
2025

Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals

NeurIPS 2025spotlight

Estimating motion primitives from video (e.g., optical flow and occlusion) is a critically important computer vision problem with many downstream applications, including controllable video generation and robotics. Current solutions are primarily supervised on synthetic data or require tuning of situ…

Cited by 0SourceScholar
2025

Taming generative video models for zero-shot optical flow extraction

NeurIPS 2025poster

Extracting optical flow from videos remains a core computer vision problem. Motivated by the recent success of large general-purpose models, we ask whether frozen self-supervised video models trained only to predict future frames can be prompted, without fine-tuning, to output flow. Prior attempts t…

Cited by 0SourceScholar
2024

Understanding Physical Dynamics with Counterfactual World Modeling

ECCV 2024poster

"The ability to understand physical dynamics is critical for agents to act in the world. Here, we use Counterfactual World Modeling (CWM) to extract vision structures for dynamics understanding. CWM uses a temporally-factored masking policy for masked prediction of video data without annotations. Th…

2023

Are These the Same Apple? Comparing Images Based on Object Intrinsics

NeurIPS 2023poster

The human visual system can effortlessly recognize an object under different extrinsic factors such as lighting, object poses, and background, yet current computer vision systems often struggle with these variations. An important step to understanding and improving artificial vision systems is to me…

2022

Break and Make: Interactive Structural Understanding Using LEGO Bricks

ECCV 2022poster

"Visual understanding of geometric structures with complex spatial relationships is a fundamental component of human intelligence. As children, we learn how to reason about structure not only from observation, but also by interacting with the world around us - by taking things apart and putting them…

2021

Contrasting Contrastive Self-Supervised Representation Learning Pipelines

ICCV 2021poster

In the past few years, we have witnessed remarkable breakthroughs in self-supervised representation learning. Despite the success and adoption of representations learned through this paradigm, much is yet to be understood about how different training methods and datasets influence performance on dow…

Cited by 63PDFcodeScholar