← Search

Vaishakh Patil

12 accepted papers

2026

DiskChunGS: Large-Scale 3D Gaussian SLAM Through Chunk-Based Memory Management

RA-L 2026

Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated impressive results for novel view synthesis with real-time rendering capabilities. However, integrating 3DGS with SLAM systems faces a fundamental scalability limitation: methods are constrained by GPU memory capacity, restricting rec

Cited by 1SourcecodeScholar
2026

MOSAIC-GS: Monocular Scene Reconstruction via Advanced Initialization for Complex Dynamic Environments

CVPR 2026

We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos using Gaussian Splatting.Monocular reconstruction is inherently ill-posed due to the lack of sufficient multiview constraints, making accurate re

Cited by 0SourceScholar
2026

ViserDex: Visual Sim-to-Real for Robust Dexterous In-hand Reorientation

RSS 2026poster

In-hand object reorientation requires precise estimation of the object pose to handle complex task dynamics. While RGB sensing offers rich semantic cues for pose tracking, existing solutions rely on multi-camera setups or costly ray tracing. We present a sim-to-real framework for monocular RGB in-ha…

Cited by 0SourceScholar
2025

FOCI: Trajectory Optimization on Gaussian Splats

IROS 2025

3D Gaussian Splatting (3DGS) has recently gained popularity as a faster alternative to Neural Radiance Fields (NeRFs) in 3D reconstruction and view synthesis methods. Leveraging the spatial information encoded in 3DGS, this work proposes FOCI (Field Overlap Collision Integral), an algorithm that is

Cited by 2SourceScholar
2024

ICGNet: A Unified Approach for Instance-Centric Grasping

ICRA 2024poster

Accurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of scene understanding: First, the robot needs to analyze the geometric properties of individual objects to find feasible…

Cited by 13SourcecodeScholar
2024

TULIP: Transformer for Upsampling of LiDAR Point Clouds

CVPR 2024poster

LiDAR Upsampling is a challenging task for the perception systems of robots and autonomous vehicles due to the sparse and irregular structure of large-scale scene contexts. Recent works propose to solve this problem by converting LiDAR data from 3D Euclidean space into an image super-resolution prob…

2024

Tag Map: A Text-Based Map for Spatial Reasoning and Navigation with Large Language Models

CoRL 2024poster

Large Language Models (LLM) have emerged as a tool for robots to generate task plans using common sense reasoning. For the LLM to generate actionable plans, scene context must be provided, often through a map. Recent works have shifted from explicit maps with fixed semantic classes to implicit open…

Cited by 3SourceScholar
2022

Lidar Line Selection with Spatially-Aware Shapley Value for Cost-Efficient Depth Completion

CoRL 2022poster

Lidar is a vital sensor for estimating the depth of a scene. Typical spinning lidars emit pulses arranged in several horizontal lines and the monetary cost of the sensor increases with the number of these lines. In this work, we present the new problem of optimizing the positioning of lidar lines to…

Cited by 2SourceScholar
2022

P3Depth: Monocular Depth Estimation With a Piecewise Planarity Prior

CVPR 2022poster

Monocular depth estimation is vital for scene understanding and downstream tasks. We focus on the supervised setup, in which ground-truth depth is available only at training time. Based on knowledge about the high regularity of real 3D scenes, we propose a method that learns to selectively leverage…

Cited by 170PDFcodeScholar
2020

Don't Forget The Past: Recurrent Depth Estimation from Monocular Video

RA-L 2020

Autonomous cars need continuously updated depth information. Thus far, depth is mostly estimated independently for a single frame at a time, even if the method starts from video input. Our method produces a time series of depth maps, which makes it an ideal candidate for online learning approaches.

Cited by 152SourceScholar