← Search

Nilesh Kulkarni

10 accepted papers

2026

Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation

CVPR 2026

Fine-tuning large-scale text-to-video diffusion models to add new generative controls, such as those over physical camera parameters (e.g., shutter speed or aperture), typically requires vast, high-fidelity datasets that are difficult to acquire. In this work, we propose a data-efficient fine-tuning

Cited by 0SourceScholar
2025

SIR-DIFF: Sparse Image Sets Restoration with Multi-View Diffusion Model

CVPR 2025poster

The computer vision community has developed numerous techniques for digitally restoring true scene information from single-view degraded photographs, an important yet extremely ill-posed task. In this work, we tackle image restoration from a different perspective by jointly denoising multiple photog…

2024

FAR: Flexible Accurate and Robust 6DoF Relative Camera Pose Estimation

CVPR 2024highlight

Estimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely methods predicting pose directly using neural networks are more robust to limited overl…

Cited by 4SourcePDFScholar
2024

NIFTY: Neural Object Interaction Fields for Guided Human Motion Synthesis

CVPR 2024poster

We address the problem of generating realistic 3D motions of humans interacting with objects in a scene. Our key idea is to create a neural interaction field attached to a specific object which outputs the distance to the valid interaction manifold given a human pose as input. This interaction field…

Cited by 44SourcePDFScholar
2023

Learning To Predict Scene-Level Implicit 3D From Posed RGBD Data

CVPR 2023poster

We introduce a method that can learn to predict scene-level implicit functions for 3D reconstruction from posed RGBD data. At test time, our system maps a previously unseen RGB image to a 3D reconstruction of a scene via implicit functions. While implicit functions for 3D reconstruction have often b…

Cited by 2SourcePDFScholar
2019

3D-RelNet: Joint Object and Relational Network for 3D Prediction

ICCV 2019poster

We propose an approach to predict the 3D shape and pose for the objects present in a scene. Existing learning based methods that pursue this goal make independent predictions per object, and do not leverage the relationships amongst them. We argue that reasoning about these relationships is crucial,…

Cited by 58PDFScholar