← Search

Yash Kant

10 accepted papers

2026

Vista4D: Video Reshooting with 4D Point Clouds

CVPR 2026

We present **Vista4D**, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the same dynamics from a different camera trajectory and viewpoint. Existing vide

Cited by 0SourcecodeScholar
2025

Pippo: High-Resolution Multi-View Humans from a Single Image

CVPR 2025highlight

We present Pippo, a generative model capable of producing 1K resolution dense turnaround videos of a person from a single casually clicked photo. Pippo is a multi-view diffusion transformer and does not require any additional inputs - e.g., a fitted parametric model or camera parameters of the input…

Cited by 1SourcePDFScholar
2025

SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation

ICLR 2025poster

Methods for image-to-video generation have achieved impressive, photo-realistic quality. However, adjusting specific elements in generated videos, such as object motion or camera movement, is often a tedious process of trial and error, e.g., involving re-generating videos with different random seed…

2025

Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior

CVPR 2025poster

We present Vid2Avatar-Pro, a method to create photorealistic and animatable 3D human avatars from monocular in-the-wild videos. Building a high-quality avatar that supports animation with diverse poses from a monocular video is challenging because the observation of pose diversity and view points is…

Cited by 0SourcePDFScholar
2024

SPAD: Spatially Aware Multi-View Diffusers

CVPR 2024poster

We present SPAD a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation we repurpose a pretrained 2D diffusion model by extending its self-attention layers with cross-view interactions and fine-tune it on a high quality subset of…

Cited by 34SourcePDFScholar
2023

Invertible Neural Skinning

CVPR 2023poster

Building animatable and editable models of clothed humans from raw 3D scans and poses is a challenging problem. Existing reposing methods suffer from the limited expressiveness of Linear Blend Skinning (LBS), require costly mesh extraction to generate each new pose, and typically do not preserve sur…

2022

Housekeep: Tidying Virtual Households Using Commonsense Reasoning

ECCV 2022poster

"We introduce Housekeep, a benchmark to evaluate commonsense reasoning in the home for embodied AI. In Housekeep, an embodied agent must tidy a house by rearranging misplaced objects without explicit instructions specifying which objects need to be rearranged. Instead, the agent must learn from and…

2022

LaTeRF: Label and Text Driven Object Radiance Fields

ECCV 2022poster

"Obtaining 3D object representations is important for creating photo-realistic simulators and collecting assets for AR/VR applications. Neural fields have shown their effectiveness in learning a continuous volumetric representation of a scene from 2D images, but acquiring object representations from…

Cited by 37SourcePDFScholar
2021

Contrast and Classify: Training Robust VQA Models

ICCV 2021poster

Recent Visual Question Answering (VQA) models have shown impressive performance on the VQA benchmark but remain sensitive to small linguistic variations in input questions. Existing approaches address this by augmenting the dataset with question paraphrases from visual question generation models or…

Cited by 35PDFcodeScholar
2020

Spatially Aware Multimodal Transformers for TextVQA

ECCV 2020poster

Textual cues are essential for everyday tasks like buying groceries and using public transport. To develop this assistive technology, we study the TextVQA task, i.e., reasoning about text in images to answer a question. Existing approaches are limited in their use of spatial relations and rely on fu…