← Search

Aayush Bansal

13 accepted papers

2026

Clothe and Pose

CVPR 2026

We introduce Clothe and Pose, an image generation and editing task that enables users to try on garments while simultaneously adopting any desired pose. Our method takes a single user image, a set of garment images, and a reference pose as input, and outputs the user wearing the target garment in th

Cited by 0SourceScholar
2025

Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image

NeurIPS 2025poster

This paper proposes Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an auto-regressive, segment-by-segment generation process, eliminating the need for resource-intensive generation and len…

Cited by 0SourcecodeScholar
2023

Ego-Humans: An Ego-Centric 3D Multi-Human Benchmark

ICCV 2023oral

We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing egocentric benchmarks either capture single subject or indoor-only scenarios, which limit the generalization of computer vision algorithms…

Cited by 39PDFScholar
2022

COAP: Compositional Articulated Occupancy of People

CVPR 2022poster

We present a novel neural implicit representation for articulated human bodies. Compared to explicit template meshes, neural implicit body representations provide an efficient mechanism for modeling interactions with the environment, which is essential for human motion reconstruction and synthesis i…

Cited by 57PDFcodeScholar
2022

KeypointNeRF: Generalizing Image-Based Volumetric Avatars Using Relative Spatial Encoding of Keypoints

ECCV 2022poster

"Image-based volumetric avatars using pixel-aligned features promise generalization to unseen poses and identities. Prior work leverages global spatial encodings and multi-view geometric consistency to reduce spatial ambiguity. However, global encodings often suffer from overfitting to the distribut…

2021

Stereo Radiance Fields (SRF): Learning View Synthesis for Sparse Views of Novel Scenes

CVPR 2021poster

Recent neural view synthesis methods have achieved impressive quality and realism, surpassing classical pipelines which rely on multi-view reconstruction. State-of-the-Art methods, such as NeRF, are designed to learn a single scene with a neural network and require dense multi-view inputs. Testing o…

Cited by 261PDFScholar
2020

4D Visualization of Dynamic Events From Unconstrained Multi-View Videos

CVPR 2020poster

We present a data-driven approach for 4D space-time visualization of dynamic events from videos captured by hand-held multiple cameras. Key to our approach is the use of self-supervised neural networks specific to the scene to compose static and dynamic aspects of an event. Though captured from disc…

Cited by 83PDFScholar