← Search

Sergey Zakharov

33 accepted papers

2026

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation

ICRA 2026poster

From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. However, the majority of current robot learning datasets and benchmarks mainly focus on stationary robot arms, and the few existing humanoid datasets are either co…

2026

PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies

RSS 2026poster

A significant challenge for robot learning research is our ability to accurately measure and compare the performance of robot policies. Benchmarking in robotics is historically challenging due to the stochasticity, reproducibility, and time-consuming nature of real-world rollouts. This challenge is …

Cited by 26SourceScholar
2026

SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes

ICML 2026spotlight

Simulation has become a key tool for training and evaluating home robots at scale, yet existing environments fail to capture the diversity and physical complexity of real indoor spaces. Current scene synthesis methods produce sparsely furnished rooms that lack the dense clutter, articulated furnitur…

Cited by 0SourceScholar
2025

OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World

ICRA 2025

We would like to estimate the pose and full shape of an object from a single observation, without assuming known 3D model or category. In this work, we propose OmniShape, the first method of its kind to enable probabilistic pose and shape estimation. OmniShape is based on the key insight that shape

Cited by 1SourceScholar
2025

Robot Learning from Any Images

CoRL 2025poster

We introduce RoLA, a framework that transforms any in‑the‑wild image into an interactive, physics‑enabled robotic environment. Unlike previous methods, RoLA operates directly on a single image without requiring additional hardware or digital assets. Our framework democratizes robotic data generatio…

Cited by 0SourcecodeScholar
2025

Steerable Scene Generation with Post Training and Inference-Time Search

CoRL 2025poster

Training robots in simulation requires diverse 3D scenes that reflect the specific challenges of downstream tasks. However, scenes that satisfy strict task requirements, such as high-clutter environments with plausible spatial arrangement, are rare and costly to curate manually. Instead, we generate…

Cited by 0SourcecodeScholar
2025

ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping

CVPR 2025poster

Robotic grasping is a cornerstone capability of embodied systems. Many methods directly output grasps from partial information without modeling the geometry of the scene, leading to suboptimal motion and even collisions. To address these issues, we introduce ZeroGrasp, a novel framework that simulta…

Cited by 0SourcePDFScholar
2024

$SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation

NeurIPS 2024poster

Incorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D learning. Equivariance serves as a valuable inductive prior,…

Cited by 1SourcePDFScholar
2024

DiffusionNOCS: Managing Symmetry and Uncertainty in Sim2Real Multi-Modal Category-level Pose Estimation

IROS 2024poster

This paper addresses the challenging problem of category-level pose estimation. Current state-of-the-art methods for this task face challenges when dealing with symmetric objects and when attempting to generalize to new environments solely through synthetic data training. In this work, we address th…

Cited by 11SourcecodeScholar
2024

FSD: Fast Self-Supervised Single RGB-D to Categorical 3D Objects

ICRA 2024poster

In this work, we address the challenging task of 3D object recognition without the reliance on real-world 3D labeled data. Our goal is to predict the 3D shape, size, and 6D pose of objects within a single RGB-D image, operating at the category level and eliminating the need for CAD models during inf…

Cited by 13SourcecodeScholar
2024

NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields

ECCV 2024poster

"Neural fields excel in computer vision and robotics due to their ability to understand the 3D visual world such as inferring semantics, geometry, and dynamics. Given the capabilities of neural fields in densely representing a 3D scene from 2D images, we ask the question: Can we scale their self-sup…

2024

View-Invariant Policy Learning via Zero-Shot Novel View Synthesis

CoRL 2024poster

Large-scale visuomotor policy learning is a promising approach toward developing generalizable manipulation systems. Yet, policies that can be deployed on diverse embodiments, environments, and observational modalities remain elusive. In this work, we investigate how knowledge from large-scale visu…

Cited by 10SourceScholar
2024

Zero-Shot Multi-Object Scene Completion

ECCV 2024poster

"We present a 3D scene completion method that recovers the complete geometry of multiple unseen objects in complex scenes from a single RGB-D image. Despite notable advancements in single-object 3D shape completion, high-quality reconstructions in highly cluttered real-world multi-object scenes rema…

Cited by 1SourcePDFScholar
2023

CARTO: Category and Joint Agnostic Reconstruction of ARTiculated Objects

CVPR 2023poster

We present CARTO, a novel approach for reconstructing multiple articulated objects from a single stereo RGB observation. We use implicit object-centric representations and learn a single geometry and articulation decoder for multiple object categories. Despite training on multiple categories, our de…

2023

DeLiRa: Self-Supervised Depth, Light, and Radiance Fields

ICCV 2023poster

Differentiable volumetric rendering is a powerful paradigm for 3D reconstruction and novel view synthesis. However, standard volume rendering approaches struggle with degenerate geometries in the case of limited viewpoint diversity, a common scenario in robotics applications. In this work, we propos…

Cited by 4PDFScholar
2023

Multi-Object Manipulation via Object-Centric Neural Scattering Functions

CVPR 2023poster

Learned visual dynamics models have proven effective for robotic manipulation tasks. Yet, it remains unclear how best to represent scenes involving multi-object interactions. Current methods decompose a scene into discrete objects, yet they struggle with precise modeling and manipulation amid challe…

Cited by 11SourcePDFScholar
2023

NeO 360: Neural Fields for Sparse View Synthesis of Outdoor Scenes

ICCV 2023poster

Recent implicit neural representations have shown great results for novel view synthesis. However, existing methods require expensive per-scene optimization from many views hence limiting their application to real-world unbounded urban settings where the objects of interest or backgrounds are observ…

Cited by 47PDFcodeScholar
2023

Neural Groundplans: Persistent Neural Scene Representations from a Single Image

ICLR 2023poster

We present a method to map 2D image observations of a scene to a persistent 3D scene representation, enabling novel view synthesis and disentangled representation of the movable and immovable components of the scene. Motivated by the bird’s-eye-view (BEV) representation commonly used in vision and r…

Cited by 13SourcePDFScholar
2023

Zero-1-to-3: Zero-shot One Image to 3D Object

ICCV 2023poster

We introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this underconstrained setting, we capitalize on the geometric priors that large-scale diffusion models learn about natural images. Our conditional di…

Cited by 1020PDFcodeScholar
2022

"ShAPO: Implicit Representations for Multi-Object Shape, Appearance, and Pose Optimization"

ECCV 2022poster

"Our method studies the complex task of object-centric 3D understanding from a single RGB-D observation. As it is an ill-posed problem, existing methods suffer from low performance for both 3D shape and 6D pose and size estimation in complex multi-object scenarios with occlusions. We present ShAPO,…

2022

Multi-Frame Self-Supervised Depth With Transformers

CVPR 2022poster

Multi-frame depth estimation improves over single-frame approaches by also leveraging geometric relationships between images via feature matching, in addition to learning appearance-based features. In this paper we revisit feature matching for self-supervised monocular depth estimation, and propose…

Cited by 109PDFScholar
2022

Photo-Realistic Neural Domain Randomization

ECCV 2022poster

"Synthetic data is a scalable alternative to manual supervision, but it requires overcoming the sim-to-real domain gap. This discrepancy between virtual and real worlds is addressed by two seemingly opposed approaches: improving the realism of simulation or foregoing realism entirely via domain rand…

Cited by 12SourcePDFScholar
2022

ROAD: Learning an Implicit Recursive Octree Auto-Decoder to Efficiently Encode 3D Shapes

CoRL 2022poster

Compact and accurate representations of 3D shapes are central to many perception and robotics tasks. State-of-the-art learning-based methods can reconstruct single objects but scale poorly to large datasets. We present a novel recursive implicit representation to efficiently and accurately encode la…

Cited by 6SourceScholar
2022

SpOT: Spatiotemporal Modeling for 3D Object Tracking

ECCV 2022poster

"3D multi-object tracking aims to uniquely and consistently identify all mobile entities through time. Despite the rich spatiotemporal information available in this setting, current 3D tracking methods primarily rely on abstracted information and limited history, e.g. single-frame object bounding bo…

Cited by 13SourcePDFScholar
2021

Single-Shot Scene Reconstruction

CoRL 2021poster

We introduce a novel scene reconstruction method to infer a fully editable and re-renderable model of a 3D road scene from a single image. We represent movable objects separately from the immovable background, and recover a full 3D model of each distinct object as well as their spatial relations in…

Cited by 18SourceScholar
2020

Autolabeling 3D Objects With Differentiable Rendering of SDF Shape Priors

CVPR 2020oral

We present an automatic annotation pipeline to recover 9D cuboids and 3D shapes from pre-trained off-the-shelf 2D detectors and sparse LIDAR data. Our autolabeling method solves an ill-posed inverse problem by considering learned shape priors and optimizing geometric and physical parameters. To addr…

Cited by 125PDFcodeScholar
2019

Seeing Beyond Appearance - Mapping Real Images into Geometrical Domains for Unsupervised CAD-based Recognition

IROS 2019poster

While convolutional neural networks are dominating the field of computer vision, one usually does not have access to the large amount of domain-relevant data needed for their training. Therefore, it has become common practice to use available synthetic samples along domain adaptation schemes to prep…

Cited by 14SourceScholar
2018

When Regression Meets Manifold Learning for Object Recognition and Pose Estimation

ICRA 2018poster

In this work, we propose a method for object recognition and pose estimation from depth images using convolutional neural networks. Previous methods addressing this problem rely on manifold learning to learn low dimensional viewpoint descriptors and employ them in a nearest neighbor search on an est…

Cited by 35SourceScholar
2017

3D object instance recognition and pose estimation using triplet loss with dynamic margin

IROS 2017poster

In this paper, we address the problem of 3D object instance recognition and pose estimation of localized objects in cluttered environments using convolutional neural networks. Inspired by the descriptor learning approach of Wohlhart et al. [1], we propose a method that introduces the dynamic margin…

Cited by 48SourceScholar