← Search

Rares Ambrus

32 accepted papers

2025

AllTracker: Efficient Dense Point Tracking at High Resolution

ICCV 2025poster

We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing point tracking methods, our approach delivers high-resolution and dense (all-pixel) correspondence fields, which can be…

2025

OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World

ICRA 2025

We would like to estimate the pose and full shape of an object from a single observation, without assuming known 3D model or category. In this work, we propose OmniShape, the first method of its kind to enable probabilistic pose and shape estimation. OmniShape is based on the key insight that shape

Cited by 1SourceScholar
2025

Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

CVPR 2025poster

Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of…

Cited by 0SourcePDFScholar
2025

ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping

CVPR 2025poster

Robotic grasping is a cornerstone capability of embodied systems. Many methods directly output grasps from partial information without modeling the geometry of the scene, leading to suboptimal motion and even collisions. To address these issues, we introduce ZeroGrasp, a novel framework that simulta…

Cited by 0SourcePDFScholar
2024

DiffusionNOCS: Managing Symmetry and Uncertainty in Sim2Real Multi-Modal Category-level Pose Estimation

IROS 2024poster

This paper addresses the challenging problem of category-level pose estimation. Current state-of-the-art methods for this task face challenges when dealing with symmetric objects and when attempting to generalize to new environments solely through synthetic data training. In this work, we address th…

Cited by 11SourcecodeScholar
2024

FSD: Fast Self-Supervised Single RGB-D to Categorical 3D Objects

ICRA 2024poster

In this work, we address the challenging task of 3D object recognition without the reliance on real-world 3D labeled data. Our goal is to predict the 3D shape, size, and 6D pose of objects within a single RGB-D image, operating at the category level and eliminating the need for CAD models during inf…

Cited by 13SourcecodeScholar
2024

NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields

ECCV 2024poster

"Neural fields excel in computer vision and robotics due to their ability to understand the 3D visual world such as inferring semantics, geometry, and dynamics. Given the capabilities of neural fields in densely representing a 3D scene from 2D images, we ask the question: Can we scale their self-sup…

2024

ShaSTA: Modeling Shape and Spatio-Temporal Affinities for 3D Multi-Object Tracking

RA-L 2024

Multi-object tracking (MOT) is a cornerstone capability of any robotic system. Tracking quality is largely dependent on the quality of input detections. In many applications, such as autonomous driving, it is preferable to over-detect objects to avoid catastrophic outcomes due to missed detections.

Cited by 44SourcecodeScholar
2024

Transcrib3D: 3D Referring Expression Resolution through Large Language Models

IROS 2024poster

If robots are to work effectively alongside people, they must be able to interpret natural language references to objects in their 3D environment. Understanding 3D referring expressions is challenging—it requires the ability to both parse the 3D structure of the scene and correctly ground free-form…

Cited by 5SourcecodeScholar
2024

Understanding Video Transformers via Universal Concept Discovery

CVPR 2024highlight

This paper studies the problem of concept-based interpretability of transformer representations for videos. Concretely we seek to explain the decision-making process of video transformers based on high-level spatiotemporal concepts that are automatically discovered. Prior research on concept-based i…

Cited by 6SourcePDFScholar
2023

DeLiRa: Self-Supervised Depth, Light, and Radiance Fields

ICCV 2023poster

Differentiable volumetric rendering is a powerful paradigm for 3D reconstruction and novel view synthesis. However, standard volume rendering approaches struggle with degenerate geometries in the case of limited viewpoint diversity, a common scenario in robotics applications. In this work, we propos…

Cited by 4PDFScholar
2023

NeO 360: Neural Fields for Sparse View Synthesis of Outdoor Scenes

ICCV 2023poster

Recent implicit neural representations have shown great results for novel view synthesis. However, existing methods require expensive per-scene optimization from many views hence limiting their application to real-world unbounded urban settings where the objects of interest or backgrounds are observ…

Cited by 47PDFcodeScholar
2023

Robust Self-Supervised Extrinsic Self-Calibration

IROS 2023poster

Autonomous vehicles and robots need to operate over a wide variety of scenarios in order to complete tasks efficiently and safely. Multi-camera self-supervised monocular depth estimation from videos is a promising way to reason about the environment, as it generates metrically scaled geometric predi…

Cited by 6SourceScholar
2023

Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?

ICRA 2023poster

Building 3D perception systems for autonomous vehicles that do not rely on high-density LiDAR is a critical research problem because of the expense of LiDAR systems compared to cameras and other sensors. Recent research has developed a variety of camera-only methods, where features are differentiabl…

Cited by 143SourceScholar
2022

Learning Optical Flow, Depth, and Scene Flow Without Real-World Labels

RA-L 2022

Self-supervised monocular depth estimation enables robots to learn 3D perception from raw video streams. This scalable approach leverages projective geometry and ego-motion to learn via view synthesis, assuming the world is mostly static. Dynamic scenes, which are common in autonomous driving and hu

Cited by 63SourceScholar
2022

Self-Supervised Camera Self-Calibration from Video

ICRA 2022poster

Camera calibration is integral to robotics and computer vision algorithms that seek to infer geometric properties of the scene from visual input streams. In practice, calibration is a laborious procedure requiring specialized data collection and careful tuning. This process must be repeated whenever…

Cited by 31SourceScholar
2021

Is Pseudo-Lidar Needed for Monocular 3D Object Detection?

ICCV 2021poster

Recent progress in 3D object detection from single images leverages monocular depth estimation as a way to produce 3D pointclouds, turning cameras into pseudo-lidar sensors. These two-stage detectors improve with the accuracy of the intermediate depth estimation network, which can itself be improved…

Cited by 388PDFcodeScholar
2021

Sparse Auxiliary Networks for Unified Monocular Depth Prediction and Completion

CVPR 2021poster

Estimating scene geometry from cost-effective sensors is key for robots. In this paper, we study the problem of predicting dense depth from a single RGB image (monodepth) with optional sparse measurements from low-cost active depth sensors. We introduce Sparse Auxiliary Networks (SAN), a new module…

Cited by 86PDFcodeScholar
2021

Warp-Refine Propagation: Semi-Supervised Auto-Labeling via Cycle-Consistency

ICCV 2021poster

Deep learning models for semantic segmentation rely on expensive, large-scale, manually annotated datasets. Labelling is a tedious process that can take hours per image. Automatically annotating video sequences by propagating sparsely labeled frames through time is a more scalable alternative. In th…

Cited by 23PDFScholar
2020

3D Packing for Self-Supervised Monocular Depth Estimation

CVPR 2020oral

Although cameras are ubiquitous, robotic platforms typically rely on active sensors like LiDAR for direct 3D perception. In this work, we propose a novel self-supervised monocular depth estimation method combining geometry with a new deep network, PackNet, learned only from unlabeled monocular video…

Cited by 881PDFcodeScholar
2020

Driving Through Ghosts: Behavioral Cloning with False Positives

IROS 2020poster

Safe autonomous driving requires robust detection of other traffic participants. However, robust does not mean perfect, and safe systems typically minimize missed detections at the expense of a higher false positive rate. This results in conservative and yet potentially dangerous behavior such as av…

Cited by 24SourceScholar
2020

Neural Outlier Rejection for Self-Supervised Keypoint Learning

ICLR 2020poster

Identifying salient points in images is a crucial component for visual odometry, Structure-from-Motion or SLAM algorithms. Recently, several learned keypoint methods have demonstrated compelling performance on challenging benchmarks. However, generating consistent and accurate training data for int…

Cited by 39SourcecodeScholar
2020

Self-Supervised 3D Keypoint Learning for Ego-Motion Estimation

CoRL 2020

Detecting and matching robust viewpoint-invariant keypoints is critical for visual SLAM and Structure-from-Motion. State-of-the-art learning-based methods generate training samples via homography adaptation to create 2D synthetic views with known keypoint matches from a single image. This approach d

2020

Semantically-Guided Representation Learning for Self-Supervised Monocular Depth

ICLR 2020poster

Self-supervised learning is showing great promise for monocular depth estimation, using geometry as the only source of supervision. Depth networks are indeed capable of learning representations that relate visual appearance to 3D properties by implicitly leveraging category-level patterns. In this w…

Cited by 285SourcecodeScholar
2019

Robust Semi-Supervised Monocular Depth Estimation with Reprojected Distances

CoRL 2019

Dense depth estimation from a single image is a key problem in computer vision, with exciting applications in a multitude of robotic tasks. Initially viewed as a direct regression problem, requiring annotated labels as supervision at training time, in the past few years a substantial amount of work

Cited by 0SourcePDFScholar
2019

Two Stream Networks for Self-Supervised Ego-Motion Estimation

CoRL 2019

Learning depth and camera ego-motion from raw unlabeled RGB video streams is seeing exciting progress through self-supervision from strong geometric cues. To leverage not only appearance but also scene geometry, we propose a novel self-supervised two-stream network using RGB and inferred depth infor

Cited by 0SourcePDFScholar
2017

Autonomous Learning of Object Models on a Mobile Robot

RA-L 2017

In this article, we present and evaluate a system, which allows a mobile robot to autonomously detect, model, and re-recognize objects in everyday environments. While other systems have demonstrated one of these elements, to our knowledge, we present the first system, which is capable of doing all o

Cited by 73SourceScholar
2017

Autonomous meshing, texturing and recognition of object models with a mobile robot

IROS 2017poster

We present a system for creating object models from RGB-D views acquired autonomously by a mobile robot. We create high-quality textured meshes of the objects by approximating the underlying geometry with a Poisson surface. Our system employs two optimization steps, first registering the views spati…

Cited by 8SourceScholar
2015

Unsupervised learning of spatial-temporal models of objects in a long-term autonomy scenario

IROS 2015poster

We present a novel method for clustering segmented dynamic parts of indoor RGB-D scenes across repeated observations by performing an analysis of their spatial-temporal distributions. We segment areas of interest in the scene using scene differencing for change detection. We extend the Meta-Room met…

Cited by 32SourceScholar
2015

Where's waldo at time t ? using spatio-temporal models for mobile robot search

ICRA 2015poster

We present a novel approach to mobile robot search for non-stationary objects in partially known environments. We formulate the search as a path planning problem in an environment where the probability of object occurrences at particular locations is a function of time. We propose to explicitly mode…

Cited by 55SourceScholar