← Search

Shalini De Mello

39 accepted papers

2025

BLADE: Single-view Body Mesh Estimation through Accurate Depth Estimation

CVPR 2025poster

Single-image human mesh recovery is a challenging task due to the ill-posed nature of simultaneous body shape, pose, and camera estimation. Existing estimators work well on images taken from afar, but they break down as the person moves close to the camera. Moreover, current methods fail to achieve…

Cited by 0SourcePDFScholar
2025

Coherent 3D Portrait Video Reconstruction via Triplane Fusion

CVPR 2025poster

Recent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user's appearance. On the…

Cited by 1SourcePDFScholar
2025

Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation

NeurIPS 2025poster

Synthetic video generation is progressing very rapidly. The latest models can produce very realistic high-resolution videos that are virtually indistinguishable from real ones. Although several video forensic detectors have been recently proposed, they often exhibit poor generalization, which limits…

Cited by 0SourceScholar
2025

SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing

CVPR 2025poster

We introduce SimAvatar, a framework designed to generate simulation-ready clothed 3D human avatars from a text prompt. Current text-driven human avatar generation methods either model hair, clothing and human body using a unified geometry or produce hair and garments that are not easily adaptable fo…

Cited by 1SourcePDFScholar
2024

3D Reconstruction with Generalizable Neural Fields using Scene Priors

ICLR 2024poster

High-fidelity 3D scene reconstruction has been substantially advanced by recent progress in neural fields. However, most existing methods train a separate network from scratch for each individual scene. This is not scalable, inefficient, and unable to yield good results given limited views. While le…

2024

A Unified Approach for Text- and Image-guided 4D Scene Generation

CVPR 2024poster

Large-scale diffusion generative models are greatly simplifying image video and 3D asset creation from user provided text prompts and images. However the challenging problem of text-to-4D dynamic 3D scene generation with diffusion guidance remains largely unexplored. We propose Dream-in-4D which fea…

Cited by 51SourcePDFScholar
2024

Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos

ECCV 2024poster

"Modern avatar generators allow anyone to synthesize photorealistic real-time talking avatars, ushering in a new era of avatar-based human communication, such as with immersive AR/VR interactions or videoconferencing with limited bandwidths. Their safe adoption, however, requires a mechanism to veri…

Cited by 3SourcePDFScholar
2024

GAvatar: Animatable 3D Gaussian Avatars with Implicit Mesh Learning

CVPR 2024highlight

Gaussian splatting has emerged as a powerful 3D representation that harnesses the advantages of both explicit (mesh) and implicit (NeRF) 3D representations. In this paper we seek to leverage Gaussian splatting to generate realistic animatable avatars from textual descriptions addressing the limitati…

Cited by 41SourcePDFScholar
2024

QUEEN: QUantized Efficient ENcoding of Dynamic Gaussians for Streaming Free-viewpoint Videos

NeurIPS 2024poster

Online free-viewpoint video (FVV) streaming is a challenging problem, which is relatively under-explored. It requires incremental on-the-fly updates to a volumetric representation, fast training and rendering to satisfy realtime constraints and a small memory footprint for efficient transmission. If…

Cited by 0SourcePDFScholar
2024

RegionGPT: Towards Region Understanding Vision Language Model

CVPR 2024poster

Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs yet they struggle with detailed regional visual understanding due to limited spatial awareness of the vision encoder and the use of coarse-grained training…

Cited by 44SourcePDFScholar
2024

What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs

CVPR 2024poster

3D-aware Generative Adversarial Networks (GANs) have shown remarkable progress in learning to generate multi-view-consistent images and 3D geometries of scenes from collections of 2D images via neural volume rendering. Yet the significant memory and computational costs of dense sampling in volume re…

Cited by 8SourcePDFScholar
2023

Affordance Diffusion: Synthesizing Hand-Object Interactions

CVPR 2023poster

Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture transfer or inserting objects into a user-specified region. In contrast, in this work…

2023

Convolutional State Space Models for Long-Range Spatiotemporal Modeling

NeurIPS 2023poster

Effectively modeling long spatiotemporal sequences is challenging due to the need to model complex spatial correlations and long-range temporal dependencies simultaneously. ConvLSTMs attempt to address this by updating tensor-valued states with recurrent neural networks, but their sequential computa…

Cited by 24SourcePDFScholar
2023

GPViT: A High Resolution Non-Hierarchical Vision Transformer with Group Propagation

ICLR 2023top-25%

We present the Group Propagation Vision Transformer (GPViT): a novel non- hierarchical (i.e. non-pyramidal) transformer model designed for general visual recognition with high-resolution features. High-resolution features (or tokens) are a natural fit for tasks that involve perceiving fine-grained d…

2023

GazeNeRF: 3D-Aware Gaze Redirection With Neural Radiance Fields

CVPR 2023poster

We propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and eye balls are separate 3D structures that move in a coordinated…

2023

Generalizable One-shot 3D Neural Head Avatar

NeurIPS 2023poster

We present a method that reconstructs and animates a 3D head avatar from a single-view portrait image. Existing methods either involve time-consuming optimization for a specific person with multiple images, or they struggle to synthesize intricate appearance details beyond the facial region. To addr…

Cited by 31SourcePDFScholar
2023

Generative Novel View Synthesis with 3D-Aware Diffusion Models

ICCV 2023oral

We present a diffusion-based model for 3D-aware generative novel view synthesis from as few as a single input image. Our model samples from the distribution of possible renderings consistent with the input and, even in the presence of ambiguity, is capable of rendering diverse and plausible novel vi…

Cited by 235PDFcodeScholar
2023

Open-Vocabulary Panoptic Segmentation With Text-to-Image Diffusion Models

CVPR 2023highlight

We present ODISE: Open-vocabulary DIffusion-based panoptic SEgmentation, which unifies pre-trained text-image diffusion and discriminative models to perform open-vocabulary panoptic segmentation. Text-to-image diffusion models have the remarkable ability to generate high-quality images with diverse…

2023

Zero-Shot Pose Transfer for Unrigged Stylized 3D Characters

CVPR 2023poster

Transferring the pose of a reference avatar to stylized 3D characters of various shapes is a fundamental task in computer graphics. Existing methods either require the stylized characters to be rigged, or they use the stylized character in the desired pose as ground truth at training. We present a z…

2022

CoordGAN: Self-Supervised Dense Correspondences Emerge From GANs

CVPR 2022poster

Recent advances show that Generative Adversarial Networks (GANs) can synthesize images with smooth variations along semantically meaningful latent directions, such as pose, expression, layout, etc. While this indicates that GANs implicitly learn pixel-level correspondences across images, few studies…

Cited by 22PDFcodeScholar
2022

Efficient Geometry-Aware 3D Generative Adversarial Networks

CVPR 2022oral

Unsupervised generation of high-quality multi-view-consistent images and 3D shapes using only collections of single-view 2D photographs has been a long-standing challenge. Existing 3D GANs are either compute-intensive or make approximations that are not 3D-consistent; the former limits quality and r…

Cited by 1564PDFcodeScholar
2022

FreeSOLO: Learning To Segment Objects Without Annotations

CVPR 2022poster

Instance segmentation is a fundamental vision task that aims to recognize and segment each object in an image. However, it requires costly annotations such as bounding boxes and segmentation masks for learning. In this work, we propose a fully unsupervised learning method that learns class-agnostic…

Cited by 136PDFcodeScholar
2022

GroupViT: Semantic Segmentation Emerges From Text Supervision

CVPR 2022poster

Grouping and recognition are important components of visual scene understanding, e.g., for object detection and semantic segmentation. With end-to-end deep learning systems, grouping of image regions usually happens implicitly via top-down supervision from pixel-level recognition labels. Instead, in…

Cited by 612PDFcodeScholar
2022

Learning Continuous Environment Fields via Implicit Functions

ICLR 2022poster

We propose a novel scene representation that encodes reaching distance -- the distance between any position in the scene to a goal along a feasible trajectory. We demonstrate that this environment field representation can directly guide the dynamic behaviors of agents in 2D mazes or 3D indoor scenes…

Cited by 12SourcePDFScholar
2021

Contrastive Syn-to-Real Generalization

ICLR 2021poster

Training on synthetic data can be beneficial for label or data-scarce scenarios. However, synthetically trained models often suffer from poor generalization in real domains due to domain gaps. In this work, we make a key observation that the diversity of the learned feature embeddings plays an impor…

2021

Learning to Track Instances without Video Annotations

CVPR 2021poster

Tracking segmentation masks of multiple instances has been intensively studied, but still faces two fundamental challenges: 1) the requirement of large-scale, frame-wise annotation, and 2) the complexity of two-stage approaches. To resolve these challenges, we introduce a novel semi-supervised frame…

Cited by 32PDFScholar
2021

Self-Supervised Object Detection via Generative Image Synthesis

ICCV 2021poster

We present SSOD -- the first end-to-end analysis-by-synthesis framework with controllable GANs for the task of self-supervised object detection. We use collections of real-world images without bounding box annotations to learn to synthesize and detect objects. We leverage controllable GANs to synthe…

Cited by 15PDFcodeScholar
2021

Weakly-Supervised Physically Unconstrained Gaze Estimation

CVPR 2021poster

A major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environments are abundantly available and can be much more easily annotated with frame-…

Cited by 45PDFcodeScholar
2020

Online Adaptation for Consistent Mesh Reconstruction in the Wild

NeurIPS 2020poster

This paper presents an algorithm to reconstruct temporally consistent 3D meshes of deformable object instances from videos in the wild. Without requiring annotations of 3D mesh, 2D keypoints, or camera pose for each video frame, we pose video-based reconstruction as a self-supervised online adaptati…

Cited by 61SourcePDFScholar
2020

Self-Learning Transformations for Improving Gaze and Head Redirection

NeurIPS 2020poster

Many computer vision tasks rely on labeled data. Rapid progress in generative modeling has led to the ability to synthesize photorealistic images. However, controlling specific aspects of the generation process such that the data can be used for supervision of downstream tasks remains challenging. I…

2020

Self-Supervised Viewpoint Learning From Image Collections

CVPR 2020poster

Training deep neural networks to estimate the viewpoint of objects requires large labeled training datasets. However, manually labeling viewpoints is notoriously hard, error-prone, and time-consuming. On the other hand, it is relatively easy to mine many unlabeled images of an object category from t…

Cited by 45PDFcodeScholar
2020

Self-supervised Single-view 3D Reconstruction via Semantic Consistency

ECCV 2020poster

We learn a self-supervised, single-view 3D reconstruction model that predicts the 3D mesh shape, texture and camera pose of a target object with a collection of 2D images and silhouettes. The proposed method does not necessitate 3D supervision, manually annotated keypoints, multi-view images of an o…

Cited by 200SourcePDFScholar
2019

Few-Shot Adaptive Gaze Estimation

ICCV 2019oral

Inter-personal anatomical differences limit the accuracy of person-independent gaze estimation networks. Yet there is a need to lower gaze errors further to enable applications requiring higher quality. Further gains can be achieved by personalizing gaze networks, ideally with few calibration sample…

Cited by 248PDFcodeScholar
2019

Joint-task Self-supervised Learning for Temporal Correspondence

NeurIPS 2019poster

This paper proposes to learn reliable dense correspondence from videos in a self-supervised manner. Our learning process integrates two highly related tasks: tracking large image regions and establishing fine-grained pixel-level associations between consecutive video frames. We exploit the synergy b…

2018

Switchable Temporal Propagation Network

ECCV 2018poster

Videos contain highly redundant information between frames. Such redundancy has been studied extensively in video compression and encoding but is less explored for more advanced video processing. In this paper, we propose a learnable unified framework for propagating a variety of visual properties o…

Cited by 49SourcePDFScholar
2017

Dynamic Facial Analysis: From Bayesian Filtering to Recurrent Neural Network

CVPR 2017poster

Facial analysis in videos, including head pose estimation and facial landmark localization, is key for many applications such as facial animation capture, human activity recognition, and human-computer interaction. In this paper, we propose to use a recurrent neural network (RNN) for joint estimatio…

Cited by 163PDFScholar
2017

Learning Affinity via Spatial Propagation Networks

NeurIPS 2017poster

In this paper, we propose a spatial propagation networks for learning affinity matrix. We show that by constructing a row/column linear propagation model, the spatially variant transformation matrix constitutes an affinity matrix that models dense, global pairwise similarities of an image. Specifica…

Cited by 338SourcePDFScholar