← Search

Goutam Bhat

16 accepted papers

2026

Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding

CVPR 2026

The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present a controllable, direct end-to-end method for upgrading a conventional video to a binocular one. Our approach, based on (conditional) latent diffusion, avoids artifacts due to explicit dep

Cited by 0SourcecodeScholar
2026

Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator

ICLR 2026oral

The rapid progress of large, pretrained models for both visual content generation and 3D reconstruction opens up new possibilities for text-to-3D generation. Intuitively, one could obtain a formidable 3D scene generator if one were able to combine the power of a modern latent text-to-video model as…

Cited by 0SourcecodeScholar
2025

LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

CVPR 2025poster

Spatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demonstrate remarkable performance on some benchmarks, they still struggle with 3D reasoning tasks like arranging objects in…

Cited by 10SourcePDFScholar
2022

Transforming Model Prediction for Tracking

CVPR 2022poster

Optimization based tracking methods have been widely successful by integrating a target model prediction module, providing effective global reasoning by minimizing an objective function. While this inductive bias integrates valuable domain knowledge, it limits the expressivity of the tracking networ…

Cited by 380PDFcodeScholar
2021

Deep Reparametrization of Multi-Frame Super-Resolution and Denoising

ICCV 2021poster

We propose a deep reparametrization of the maximum a posteriori formulation commonly employed in multi-frame image restoration tasks. Our approach is derived by introducing a learned error metric and a latent representation of the target image, which transforms the MAP objective to a deep feature sp…

Cited by 74PDFScholar
2021

Generating Masks From Boxes by Mining Spatio-Temporal Consistencies in Videos

ICCV 2021poster

Segmenting objects in videos is a fundamental computer vision task. The current deep learning based paradigm offers a powerful, but data-hungry solution. However, current datasets are limited by the cost and human effort of annotating object masks in videos. This effectively limits the performance a…

Cited by 23PDFcodeScholar
2020

Energy-Based Models for Deep Probabilistic Regression

ECCV 2020poster

While deep learning-based classification is generally tackled using standardized approaches, a wide variety of techniques are employed for regression. In computer vision, one particularly popular such technique is that of confidence-based regression, which entails predicting a confidence value for e…

2020

Know Your Surroundings: Exploiting Scene Information for Object Tracking

ECCV 2020poster

Current state-of-the-art trackers rely only on a target appearance model in order to localize the object in each frame. Such approaches are however prone to fail in case of e.g. fast appearance changes or presence of distractor objects, where a target appearance model alone is insufficient for robus…

2020

Learning What to Learn for Video Object Segmentation

ECCV 2020poster

Video object segmentation (VOS) is a highly challenging problem, since the target object is only defined by a first-frame reference mask during inference. The problem of how to capture and utilize this limited information to accurately segment the target remains a fundamental research question. We a…

2019

ATOM: Accurate Tracking by Overlap Maximization

CVPR 2019oral

While recent years have witnessed astonishing improvements in visual tracking robustness, the advancements in tracking accuracy have been limited. As the focus has been directed towards the development of powerful classifiers, the problem of accurate target state estimation has been largely overlook…

Cited by 1659PDFcodeScholar
2018

Unveiling the Power of Deep Tracking

ECCV 2018poster

In the field of generic object tracking numerous attempts have been made to exploit deep features. Despite all expectations, deep trackers are yet to reach an outstanding level of performance compared to methods solely based on handcrafted features. In this paper, we investigate this key issue and p…

Cited by 605SourcePDFScholar