← Search

Srinath Sridhar

35 accepted papers

2026

DreamControl: Human-Inspired Whole-Body Humanoid Control for Scene Interaction Via Guided Diffusion

ICRA 2026poster

We introduce DreamControl, a novel methodology for learning autonomous whole-body humanoid skills. DreamControl leverages the strengths of diffusion models and Reinforcement Learning (RL): our core innovation is the use of a diffusion prior trained on human motion data, which subsequently guides an …

2026

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens

CVPR 2026

Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language generation and understanding remains largely underexplored. Existing approaches often fine-tune large language models (LLM

Cited by 0SourcecodeScholar
2026

PackUV: Packed Gaussian UV Maps for 4D Volumetric Video

CVPR 2026

Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality reconstruction but break down on long sequences, temporal inconsistency, and fail under large motions and disocclusions. M

Cited by 0SourceScholar
2026

SLoFT: End-To-End Semantic Localization with Floorplan and Transformer

ICRA 2026poster

Visual localization is critical for AR navigation, AI-driven audio guidance, and mobile robot localization. How- ever, traditional SLAM methods that rely on pre-built 3D maps suffer from high costs, privacy concerns, and sensitivity to environmental changes. Recent floorplan-based localization metho…

Cited by 0Scholar
2026

Turbo-GS: Accelerating 3D Gaussian Fitting for High-Resolution Radiance Fields

CVPR 2026

Novel-view synthesis plays a crucial role in computer vision with applications in 3D reconstruction, mixed reality, and robotics. Recent approaches, such as 3D Gaussian Splatting (3DGS), have emerged as state-of-the-art solutions, offering high-quality novel view synthesis in real time. However, tra

Cited by 0SourcecodeScholar
2025

Da-Vil: Adaptive Dual-Arm Manipulation with Reinforcement Learning and Variable Impedance Control

ICRA 2025

Dual-arm manipulation is an area of growing interest in the robotics community. Enabling robots to perform tasks that require the coordinated use of two arms, is essential for complex manipulation tasks such as handling large objects, assembling components, and performing human-like interactions. Ho

Cited by 13SourcecodeScholar
2025

FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation

CVPR 2025highlight

Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale domain-specific diffusion model for synthesizing single and dual hand…

Cited by 0SourcePDFScholar
2025

GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities

CVPR 2025highlight

Understanding bimanual human hand activities is a critical problem in AI and robotics. We cannot build large models of bimanual activities because existing datasets lack the scale, coverage of diverse hand activities, and detailed annotations. We introduce GigaHands, a massive annotated dataset capt…

Cited by 3SourcePDFScholar
2025

InteractAvatar: Modeling Hand-Face Interaction in Photorealistic Avatars with Deformable Gaussians

ICCV 2025poster

With the rising interest from the community in digital avatars coupled with the importance of expressions and gestures in communication, modeling natural avatar behavior remains an important challenge across many industries such as teleconferencing, gaming, and AR/VR. Human hands are the primary too…

Cited by 0SourcePDFScholar
2025

UVGS: Reimagining Unstructured 3D Gaussian Splatting using UV Mapping

CVPR 2025poster

3D Gaussian Splatting (3DGS) has demonstrated superior quality in modeling 3D objects and scenes. However, generating 3DGS remains challenging due to their discrete, unstructured, and permutation-invariant nature. In this work, we present a simple yet effective method to overcome these challenges. W…

Cited by 2SourcePDFScholar
2025

V-HOP: Visuo-Haptic 6D Object Pose Tracking

RSS 2025poster

Humans naturally integrate vision and haptics for robust object perception during manipulation; losing either modality significantly degrades performance. Inspired by this multisensory integration, prior pose estimation research has attempted to combine visual and haptic/tactile feedback. While thes…

Cited by 2PDFScholar
2024

AnyHome: Open-Vocabulary Large-Scale Indoor Scene Generation with First-Person View Exploration

ECCV 2024poster

"Inspired by cognitive theories, we introduce , a framework that translates any text into well-structured and textured indoor scenes at a house-scale. By prompting Large Language Models (LLMs) with designed templates, our approach converts provided textual narratives into amodal structured represent…

Cited by 0SourcePDFScholar
2024

Constrained 6-DoF Grasp Generation on Complex Shapes for Improved Dual-Arm Manipulation

IROS 2024poster

Efficiently generating grasp poses tailored to specific regions of an object is vital for various robotic manipulation tasks, especially in a dual-arm setup. This scenario presents a significant challenge due to the complex geometries involved, requiring a deep understanding of the local geometry to…

Cited by 6SourcecodeScholar
2024

DiVa-360: The Dynamic Visual Dataset for Immersive Neural Fields

CVPR 2024highlight

Advances in neural fields are enabling high-fidelity capture of the shape and appearance of dynamic 3D scenes. However their capabilities lag behind those offered by conventional representations such as 2D videos because of algorithmic challenges and the lack of large-scale multi-view real-world dat…

Cited by 6SourcePDFScholar
2024

MANUS: Markerless Grasp Capture using Articulated 3D Gaussians

CVPR 2024poster

Understanding how we grasp objects with our hands has important applications in areas like robotics and mixed reality. However this challenging problem requires accurate modeling of the contact between hands and objects.To capture grasps existing methods use skeletons meshes or parametric models tha…

Cited by 12SourcePDFScholar
2023

CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes From Natural Language

CVPR 2023poster

Recent works have demonstrated that natural language can be used to generate and edit 3D shapes. However, these methods generate shapes with limited fidelity and diversity. We introduce CLIP-Sculptor, a method to address these constraints by producing high-fidelity and diverse 3D shapes without the…

Cited by 54SourcePDFScholar
2023

Canonical Fields: Self-Supervised Learning of Pose-Canonicalized Neural Fields

CVPR 2023highlight

Coordinate-based implicit neural networks, or neural fields, have emerged as useful representations of shape and appearance in 3D computer vision. Despite advances however, it remains challenging to build neural fields for categories of objects without datasets like ShapeNet that provide "canonicali…

2023

HyP-NeRF: Learning Improved NeRF Priors using a HyperNetwork

NeurIPS 2023poster

Neural Radiance Fields (NeRF) have become an increasingly popular representation to capture high-quality appearance and shape of scenes and objects. However, learning generalizable NeRF priors over categories of scenes or objects has been challenging due to the high dimensionality of network weight…

Cited by 13SourcePDFScholar
2023

LEGO-Net: Learning Regular Rearrangements of Objects in Rooms

CVPR 2023poster

Humans universally dislike the task of cleaning up a messy room. If machines were to help us with this task, they must understand human criteria for regular arrangements, such as several types of symmetry, co-linearity or co-circularity, spacing uniformity in linear or circular patterns, and further…

Cited by 62SourcePDFScholar
2023

SCARP: 3D Shape Completion in ARbitrary Poses for Improved Grasping

ICRA 2023poster

Recovering full 3D shapes from partial observations is a challenging task that has been extensively addressed in the computer vision community. Many deep learning methods tackle this problem by training 3D shape generation networks to learn a prior over the full 3D shapes. In this training regime, t…

Cited by 14SourcecodeScholar
2023

Semantic Attention Flow Fields for Monocular Dynamic Scene Decomposition

ICCV 2023poster

From video, we reconstruct a neural volume that captures time-varying color, density, scene flow, semantics, and attention information. The semantics and attention let us identify salient foreground objects separately from the background across spacetime. To mitigate low resolution semantic and atte…

Cited by 15PDFScholar
2023

Strata-NeRF : Neural Radiance Fields for Stratified Scenes

ICCV 2023poster

Neural Radiance Fields (NeRF) approaches learn the underlying 3D representation of a scene and generate photo-realistic novel views with high fidelity. However, most proposed settings concentrate on 3D modelling a single object or a single level of a scene. However, in the real world, a person captu…

Cited by 4PDFScholar
2022

ConDor: Self-Supervised Canonicalization of 3D Pose for Partial Shapes

CVPR 2022poster

Progress in 3D object understanding has relied on manually "canonicalized" shape datasets that contain instances with consistent position and orientation (3D pose). This has made it hard to generalize these methods to in-the-wild shapes, e.g., from internet model collections or depth sensors. ConDor…

Cited by 42PDFcodeScholar
2022

ShapeCrafter: A Recursive Text-Conditioned 3D Shape Generation Model

NeurIPS 2022accept

We present ShapeCrafter, a neural network for recursive text-conditioned 3D shape generation. Existing methods to generate text-conditioned 3D shapes consume an entire text prompt to generate a 3D shape in a single step. However, humans tend to describe shapes recursively---we may start with an init…

2021

DRACO: Weakly Supervised Dense Reconstruction And Canonicalization of Objects

ICRA 2021poster

We present DRACO, a method for Dense Reconstruction And Canonicalization of Object shape from one or more RGB images. Canonical shape reconstruction— estimating 3D object shape in a coordinate space canonicalized for scale, rotation, and translation parameters—is an emerging paradigm that holds prom…

Cited by 6SourcecodeScholar
2021

HuMoR: 3D Human Motion Model for Robust Pose Estimation

ICCV 2021poster

We introduce HuMoR: a 3D Human Motion Model for Robust Estimation of temporal pose and shape. Though substantial progress has been made in estimating 3D human motion and shape from dynamic observations, recovering plausible pose sequences in the presence of noise and occlusions remains a challenge.…

Cited by 354PDFcodeScholar
2020

CaSPR: Learning Canonical Spatiotemporal Point Cloud Representations

NeurIPS 2020spotlight

We propose CaSPR, a method to learn object-centric Canonical Spatiotemporal Point Cloud Representations of dynamically moving or evolving objects. Our goal is to enable information aggregation over time and the interrogation of object state at any spatiotemporal neighborhood in the past, observed or…

2020

Pix2Surf: Learning Parametric 3D Surface Models of Objects from Images

ECCV 2020poster

We investigate the problem of learning to generate 3D parametric surface representations for novel object instances, as seen from one or more views. Previous work on learning shape reconstruction from multiple views uses discrete representations such as point clouds or voxels, while continuous surfa…

Cited by 42SourcePDFScholar
2019

Multiview Aggregation for Learning Category-Specific Shape Reconstruction

NeurIPS 2019poster

We investigate the problem of learning category-specific 3D shape reconstruction from a variable number of RGB views of previously unobserved object instances. Most approaches for multiview shape reconstruction operate on sparse shape representations, or assume a fixed number of views. We present a…

2019

Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation

CVPR 2019oral

The goal of this paper is to estimate the 6D pose and dimensions of unseen object instances in an RGB-D image. Contrary to "instance-level" 6D pose estimation tasks, our problem assumes that no exact object CAD models are available during either training or testing time. To handle different and unse…

Cited by 870PDFcodeScholar
2018

GANerated Hands for Real-Time 3D Hand Tracking From Monocular RGB

CVPR 2018poster

We address the highly challenging problem of real-time 3D hand tracking based on a monocular RGB-only sequence. Our tracking method combines a convolutional neural network with a kinematic 3D hand model, such that it generalizes well to unseen data, is robust to occlusions and varying camera viewpoi…

Cited by 669SourcePDFScholar
2017

Real-Time Hand Tracking Under Occlusion From an Egocentric RGB-D Sensor

ICCV 2017poster

We present an approach for real-time, robust, and accurate hand pose estimation from moving egocentric RGB-D cameras in cluttered real environments. Existing methods typically fail for hand-object interactions in cluttered scenes imaged from egocentric viewpoints, common for virtual or augmented rea…

Cited by 409PDFScholar
2015

Fast and Robust Hand Tracking Using Detection-Guided Optimization

CVPR 2015poster

Markerless tracking of hands and fingers is a promising enabler for human-computer interaction. However, adoption has been limited because of tracking inaccuracies, incomplete coverage of motions, low framerate, complex camera setups, and high computational requirements. In this paper, we present a…

Cited by 298SourcePDFScholar