← Search

Achuta Kadambi

25 accepted papers

2026

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

CVPR 2026

Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their inability to capture fine-grained 3D geometry and spatial relationships. While recent efforts have introduced multi-view ge

Cited by 0SourceScholar
2025

4K4DGen: Panoramic 4D Generation at 4K Resolution

ICLR 2025spotlight

The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques either focus solely on dynamic objects or perform outpainting from a single perspe…

2025

Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields

CVPR 2025poster

Recent advancements in 2D and multimodal models have achieved remarkable success by leveraging large-scale training on extensive datasets. However, extending these achievements to enable free-form interactions and high-level semantic operations with complex 3D/4D scenes remains challenging. This dif…

Cited by 1SourcePDFScholar
2025

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

ICCV 2025poster

Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts--abili…

Cited by 0SourcePDFScholar
2024

All-day Depth Completion

IROS 2024poster

We propose a method for depth estimation under different illumination conditions, i.e., day and night time. As photometry is uninformative in regions under low-illumination, we tackle the problem through a multi-sensor fusion approach, where we take as input an additional synchronized sparse point c…

Cited by 3SourcecodeScholar
2024

DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting

ECCV 2024poster

"The increasing demand for virtual reality applications has highlighted the significance of crafting immersive 3D assets. We present a text-to-3D 360◦ scene generation pipeline that facilitates the creation of comprehensive 360◦ scenes for in-the-wild environments in a matter of minutes. Our approac…

2024

Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields

CVPR 2024highlight

3D scene representations have gained immense popularity in recent years. Methods that use Neural Radiance fields are versatile for traditional tasks such as novel view synthesis. In recent times some work has emerged that aims to extend the functionality of NeRF beyond view synthesis for semanticall…

2024

Implicit Neural Models to Extract Heart Rate from Video

ECCV 2024poster

"Scene representation networks, or implicit neural representations (INR) have seen a range of success in numerous image and video applications. However, being universal function fitters, they fit all variations in videos without any selectivity. This is particularly a problem for tasks such as remot…

2024

Large Spatial Model: End-to-end Unposed Images to Semantic 3D

NeurIPS 2024poster

Reconstructing and understanding 3D structures from a limited number of images is a classical problem in computer vision. Traditional approaches typically decompose this task into multiple subtasks, involving several stages of complex mappings between different data representations. For example, den…

2024

MVDD: Multi-View Depth Diffusion Models

ECCV 2024poster

"Denoising diffusion models have demonstrated outstanding results in 2D image generation, yet it remains a challenge to replicate its success in 3D shape generation. In this paper, we propose leveraging multi-view depth, which represents complex 3D shapes in a 2D data format that is easy to denoise.…

Cited by 4SourcePDFScholar
2023

ALTO: Alternating Latent Topologies for Implicit 3D Reconstruction

CVPR 2023poster

This work introduces alternating latent topologies (ALTO) for high-fidelity reconstruction of implicit 3D surfaces from noisy point clouds. Previous work identifies that the spatial arrangement of latent encodings is important to recover detail. One school of thought is to encode a latent vector for…

Cited by 35SourcePDFScholar
2023

Depth Estimation From Camera Image and mmWave Radar Point Cloud

CVPR 2023poster

We present a method for inferring dense depth from a camera image and a sparse noisy radar point cloud. We first describe the mechanics behind mmWave radar point cloud formation and the challenges that it poses, i.e. ambiguous elevation and noisy depth and azimuth components that yields incorrect po…

Cited by 53SourcePDFScholar
2023

WeatherStream: Light Transport Automation of Single Image Deweathering

CVPR 2023poster

Today single image deweathering is arguably more sensitive to the dataset type, rather than the model. We introduce WeatherStream, an automatic pipeline capturing all real-world weather effects (rain, snow, and rain fog degradations), along with their clean image pairs. Previous state-of-the-art met…

Cited by 23SourcePDFScholar
2022

MIME: Minority Inclusion for Majority Group Enhancement of AI Performance

ECCV 2022poster

"Several papers have rightly included minority groups in artificial intelligence (AI) training data to improve test inference for minority groups and/or society-at-large. A society-at-large consists of both minority and majority stakeholders. A common misconception is that minority inclusion does no…

Cited by 6SourcePDFScholar
2022

Not Just Streaks: Towards Ground Truth for Single Image Deraining

ECCV 2022poster

"We propose a large-scale dataset of real-world rainy and clean image pairs and a method to remove degradations, induced by rain streaks and rain accumulation, from the image. As there exists no real-world dataset for deraining, current state-of-the-art methods rely on synthetic data and thus are li…

2022

Synthetic Generation of Face Videos With Plethysmograph Physiology

CVPR 2022poster

Accelerated by telemedicine, advances in Remote Photoplethysmography (rPPG) are beginning to offer a viable path toward non-contact physiological measurement. Unfortunately, the datasets for rPPG are limited as they require videos of the human face paired with ground-truth, synchronized heart rate d…

Cited by 48PDFScholar
2021

Towards Rotation Invariance in Object Detection

ICCV 2021poster

Rotation augmentations generally improve a model's invariance/equivariance to rotation - except in object detection. In object detection the shape is not known, therefore rotation creates a label ambiguity. We show that the de-facto method for bounding box label rotation, the Largest Box Method, cre…

Cited by 18PDFcodeScholar
2020

Deep Polarization Cues for Transparent Object Segmentation

CVPR 2020oral

Segmentation of transparent objects is a hard, open problem in computer vision. Transparent objects lack texture of their own, adopting instead the texture of scene background. This paper reframes the problem of transparent object segmentation into the realm of light polarization, i.e., the rotation…

Cited by 169PDFcodeScholar
2020

Deep Shape from Polarization

ECCV 2020poster

This paper makes a first attempt to bring the Shape from Polarization (SfP) problem to the realm of deep learning. The previous state-of-the-art methods for SfP have been purely physics-based. We see value in these principled models, and blend these physical models as priors into a neural network ar…

2016

Macroscopic Interferometry: Rethinking Depth Estimation With Frequency-Domain Time-Of-Flight

CVPR 2016oral

A form of meter-scale, macroscopic interferometry is proposed using conventional time-of-flight (ToF) sensors. Today, ToF sensors use phase-based sampling, where the phase delay between emitted and received, high-frequency signals encodes distance. This paper examines an alternative ToF architecture…

Cited by 45PDFScholar
2015

A Light Transport Model for Mitigating Multipath Interference in Time-of-Flight Sensors

CVPR 2015poster

Continuous-wave Time-of-flight (TOF) range imaging has become a commercially viable technology with many applications in computer vision and graphics. However, the depth images obtained from TOF cameras contain scene dependent errors due to multipath interference (MPI). Specifically, MPI occurs when…

Cited by 105SourcePDFScholar