← Search

Stefan Stojanov

14 accepted papers

2026

Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

CVPR 2026

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settin

Cited by 0SourceScholar
2026

Physical Object Understanding with a Physically Controllable World Model

CVPR 2026

A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solving these tasks requires world models capable of inferring distributional states of the world from partial observations -

Cited by 0SourceScholar
2025

Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals

NeurIPS 2025spotlight

Estimating motion primitives from video (e.g., optical flow and occlusion) is a critically important computer vision problem with many downstream applications, including controllable video generation and robotics. Current solutions are primarily supervised on synthetic data or require tuning of situ…

Cited by 0SourceScholar
2025

Taming generative video models for zero-shot optical flow extraction

NeurIPS 2025poster

Extracting optical flow from videos remains a core computer vision problem. Motivated by the recent success of large general-purpose models, we ask whether frozen self-supervised video models trained only to predict future frames can be prompted, without fine-tuning, to output flow. Prior attempts t…

Cited by 0SourceScholar
2025

Weakly-Supervised Learning of Dense Functional Correspondences

ICCV 2025poster

Establishing dense correspondences across image pairs is essential for tasks such as shape reconstruction and robot manipulation. In the challenging setting of matching across different categories, the function of an object, i.e., the effect that an object can cause on other objects, can guide how c…

Cited by 0SourcePDFScholar
2024

3x2: 3D Object Part Segmentation by 2D Semantic Correspondences

ECCV 2024poster

"3D object part segmentation is essential in computer vision applications. While substantial progress has been made in 2D object part segmentation, the 3D counterpart has received less attention, in part due to the scarcity of annotated 3D datasets, which are expensive to collect. In this work, we p…

2024

ZeroShape: Regression-based Zero-shot Shape Reconstruction

CVPR 2024poster

We study the problem of single-image zero-shot 3D shape reconstruction. Recent works learn zero-shot shape reconstruction through generative modeling of 3D assets but these models are computationally expensive at train and inference time. In contrast the traditional approach to this problem is regre…

2023

Low-shot Object Learning with Mutual Exclusivity Bias

NeurIPS 2023poster

This paper introduces Low-shot Object Learning with Mutual Exclusivity Bias (LSME), the first computational framing of mutual exclusivity bias, a phenomenon commonly observed in infants during word learning. We provide a novel dataset, comprehensive baselines, and a SOTA method to enable the ML comm…

2023

ShapeClipper: Scalable 3D Shape Learning From Single-View Images via Geometric and CLIP-Based Consistency

CVPR 2023poster

We present ShapeClipper, a novel method that reconstructs 3D object shapes from real-world single-view RGB images. Instead of relying on laborious 3D, multi-view or camera pose annotation, ShapeClipper learns shape reconstruction from a set of single-view segmented images. The key idea is to facilit…

Cited by 22SourcePDFScholar
2022

Learning Dense Object Descriptors from Multiple Views for Low-shot Category Generalization

NeurIPS 2022accept

A hallmark of the deep learning era for computer vision is the successful use of large-scale labeled datasets to train feature representations. This has been done for tasks ranging from object recognition and semantic segmentation to optical flow estimation and novel view synthesis of 3D scenes. In…

2022

Planes vs. Chairs: Category-Guided 3D Shape Learning without Any 3D Cues

ECCV 2022poster

"We present a novel 3D shape reconstruction method which learns to predict an implicit 3D shape representation from a single RGB image. Our approach uses a set of single-view images of multiple object categories without viewpoint annotation, forcing the model to learn across multiple object categori…

Cited by 16SourcePDFScholar
2019

Incremental Object Learning From Contiguous Views

CVPR 2019oral

In this work, we present CRIB (Continual Recognition Inspired by Babies), a synthetic incremental object learning environment that can produce data that models visual imagery produced by object exploration in early infancy. CRIB is coupled with a new 3D object dataset, Toys-200, that contains 200 un…

Cited by 53PDFScholar
2019

Unsupervised 3D Pose Estimation With Geometric Self-Supervision

CVPR 2019poster

We present an unsupervised learning approach to re- cover 3D human pose from 2D skeletal joints extracted from a single image. Our method does not require any multi- view image data, 3D skeletons, correspondences between 2D-3D points, or use previously learned 3D priors during training. A lifting ne…

Cited by 249PDFScholar