← Search

Suryansh Kumar

19 accepted papers

2024

ICGNet: A Unified Approach for Instance-Centric Grasping

ICRA 2024poster

Accurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of scene understanding: First, the robot needs to analyze the geometric properties of individual objects to find feasible…

Cited by 13SourcecodeScholar
2024

Stereo Risk: A Continuous Modeling Approach to Stereo Matching

ICML 2024oral

We introduce Stereo Risk, a new deep-learning approach to solve the classical stereo-matching problem in computer vision. As it is well-known that stereo matching boils down to a per-pixel disparity estimation problem, the popular state-of-the-art stereo-matching approaches widely rely on regressing…

Cited by 8SourcePDFScholar
2023

How To Not Train Your Dragon: Training-free Embodied Object Goal Navigation with Semantic Frontiers

RSS 2023poster

Object goal navigation is an important problem in Embodied AI that involves guiding the agent to navigate to an instance of the object category in an unknown environment---typically an indoor scene. Unfortunately, current state-of-the-art methods for this problem rely heavily on data-driven approach…

Cited by 54SourcePDFScholar
2023

Single Image Depth Prediction Made Better: A Multivariate Gaussian Take

CVPR 2023poster

Neural-network-based single image depth prediction (SIDP) is a challenging task where the goal is to predict the scene's per-pixel depth at test time. Since the problem, by definition, is ill-posed, the fundamental goal is to come up with an approach that can reliably model the scene depth from a se…

Cited by 26SourcePDFScholar
2023

VA-DepthNet: A Variational Approach to Single Image Depth Prediction

ICLR 2023top-25%

We introduce VA-DepthNet, a simple, effective, and accurate deep neural network approach for the single-image depth prediction (SIDP) problem. The proposed approach advocates using classical first-order variational constraints for this problem. While state-of-the-art deep neural network methods for…

2022

A Real-Time Online Learning Framework for Joint 3D Reconstruction and Semantic Segmentation of Indoor Scenes

RA-L 2022

This letter presents a real-time online vision framework to jointly recover an indoor scene’s 3D structure and semantic label. Given noisy depth maps, a camera trajectory, and 2D semantic labels at train time, the proposed deep neural network based approach learns to fuse the depth over frames with

Cited by 26SourcecodeScholar
2022

CC-3DT: Panoramic 3D Object Tracking via Cross-Camera Fusion

CoRL 2022poster

To track the 3D locations and trajectories of the other traffic participants at any given time, modern autonomous vehicles are equipped with multiple cameras that cover the vehicle's full surroundings. Yet, camera-based 3D object tracking methods prioritize optimizing the single-camera setup and res…

Cited by 32SourceScholar
2022

Generative Flows With Invertible Attentions

CVPR 2022poster

Flow-based generative models have shown an excellent ability to explicitly learn the probability density function of data via a sequence of invertible transformations. Yet, learning attentions in generative flows remains understudied, while it has made breakthroughs in other domains. To fill the gap…

Cited by 16PDFcodeScholar
2022

Learning Online Multi-sensor Depth Fusion

ECCV 2022poster

"Many hand-held or mixed reality devices are used with a single sensor for 3D reconstruction, although they often comprise multiple sensors. Multi-sensor depth fusion is able to substantially improve the robustness and accuracy of 3D reconstruction methods, but existing techniques are not robust eno…

2022

Uncertainty Guided Policy for Active Robotic 3D Reconstruction Using Neural Radiance Fields

RA-L 2022

In this letter, we tackle the problem of active robotic 3D reconstruction of an object. In particular, we study how a mobile robot with an arm-held camera can select a favorable number of views to recover an object's 3D shape efficiently. Contrary to the existing solution to this problem, we leverag

Cited by 100SourceScholar
2022

Uncertainty-Aware Deep Multi-View Photometric Stereo

CVPR 2022poster

This paper presents a simple and effective solution to the longstanding classical multi-view photometric stereo (MVPS) problem. It is well-known that photometric stereo (PS) is excellent at recovering high-frequency surface details, whereas multi-view stereo (MVS) can help remove the low-frequency d…

Cited by 45PDFScholar
2021

Neural Architecture Search of SPD Manifold Networks

IJCAI 2021poster

In this paper, we propose a new neural architecture search (NAS) problem of Symmetric Positive Definite (SPD) manifold networks, aiming to automate the design of SPD neural architectures. To address this problem, we first introduce a geometrically rich and diverse SPD neural architecture search spac…

2021

Uncalibrated Neural Inverse Rendering for Photometric Stereo of General Surfaces

CVPR 2021poster

This paper presents an uncalibrated deep neural network framework for the photometric stereo problem. For training models to solve the problem, existing neural network-based methods either require exact light directions or ground-truth surface normals of the object or both. However, in practice, it…

Cited by 64PDFScholar
2018

Scalable Dense Non-Rigid Structure-From-Motion: A Grassmannian Perspective

CVPR 2018poster

This paper addresses the task of dense non-rigid structure-from-motion (NRSfM) using multiple images. State-of-the-art methods to this problem are often hurdled by scalability, expensive computations, and noisy measurements. Further, recent methods to NRSfM usually either assume a small number of sp…

Cited by 58SourcePDFScholar
2017

Monocular Dense 3D Reconstruction of a Complex Dynamic Scene From Two Perspective Frames

ICCV 2017poster

This paper proposes a new approach for monocular dense 3D reconstruction of a complex dynamic scene from two perspective frames. By applying superpixel oversegmentation to the image, we model a generically dynamic (hence non-rigid) scene with a piecewise planar and rigid approximation. In this way,…

Cited by 74PDFScholar