← Search

Anurag Ranjan

19 accepted papers

2025

INRFlow: Flow Matching for INRs in Ambient Space

ICML 2025poster

Flow matching models have emerged as a powerful method for generative modeling on domains like images or videos, and even on irregular or unstructured data like 3D point clouds or even protein structures. These models are commonly trained in two stages: first, a data compressor is trained, and in a…

Cited by 0SourcePDFScholar
2024

HUGS: Human Gaussian Splats

CVPR 2024poster

Recent advances in neural rendering have improved both training and rendering times by orders of magnitude. While these methods demonstrate state-of-the-art quality and speed they are designed for photogrammetry of static scenes and do not generalize well to freely moving humans in the environment.…

2024

Probabilistic Speech-Driven 3D Facial Motion Synthesis: New Benchmarks Methods and Applications

CVPR 2024poster

We consider the task of animating 3D facial geometry from speech signal. Existing works are primarily deterministic focusing on learning a one-to-one mapping from speech signal to 3D face meshes on small datasets with limited speakers. While these models can achieve high-quality lip articulation for…

Cited by 14SourcePDFScholar
2023

FastViT: A Fast Hybrid Vision Transformer Using Structural Reparameterization

ICCV 2023poster

The recent amalgamation of transformer and convolutional designs has led to steady improvements in accuracy and efficiency of the models. In this work, we introduce FastViT, a hybrid vision transformer architecture that obtains the state-of-the-art latency-accuracy trade-off. To this end, we intro…

Cited by 231PDFcodeScholar
2023

FineRecon: Depth-aware Feed-forward Network for Detailed 3D Reconstruction

ICCV 2023poster

Recent works on 3D reconstruction from posed images have demonstrated that direct inference of scene-level 3D geometry without test-time optimization is feasible using deep neural networks, showing remarkable promise and high efficiency. However, the reconstructed geometry, typically represented as…

Cited by 27PDFcodeScholar
2023

MobileOne: An Improved One Millisecond Mobile Backbone

CVPR 2023poster

Efficient neural network backbones for mobile devices are often optimized for metrics such as FLOPs or parameter count. However, these metrics may not correlate well with latency of the network when deployed on a mobile device. Therefore, we perform extensive analysis of different metrics by deployi…

2023

Naturalistic Head Motion Generation from Speech

ICASSP 2023accepted

Synthesizing natural head motion to accompany speech for an embodied conversational agent is necessary for pro-viding a rich interactive experience. Most prior works assess the quality of generated head motion by comparing them against a single ground-truth using an objective metric. Yet there are m…

Cited by 0SourceScholar
2023

Pointersect: Neural Rendering With Cloud-Ray Intersection

CVPR 2023poster

We propose a novel method that renders point clouds as if they are surfaces. The proposed method is differentiable and requires no scene-specific optimization. This unique capability enables, out-of-the-box, surface normal estimation, rendering room-scale point clouds, inverse rendering, and ray tra…

Cited by 20SourcePDFScholar
2022

NeuMan: Neural Human Radiance Field from a Single Video

ECCV 2022poster

"Photorealistic rendering and reposing of humans is important for enabling augmented reality experiences. We propose a novel framework to reconstruct the human and the scene that can be rendered with novel human poses and views from just a single in-the-wild video. Given a video captured by a moving…

2022

SPIN: An Empirical Evaluation on Sharing Parameters of Isotropic Networks

ECCV 2022poster

"Recent isotropic networks, such as ConvMixer and Vision Transformers, have found significant success across visual recognition tasks, matching or outperforming non-isotropic Convolutional Neural Networks. Isotropic architectures are particularly well-suited to cross-layer weight sharing, an effecti…

2021

Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding

ICCV 2021poster

For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images. We address this challenge by introducing Hypersim, a photorealistic synthetic dataset for holistic indoor scene understanding. To create our dataset, we leverage a…

Cited by 395PDFcodeScholar
2020

Learning to Dress 3D People in Generative Clothing

CVPR 2020poster

Three-dimensional human body models are widely used in the analysis of human pose and motion. Existing models, however, are learned from minimally-clothed 3D scans and thus do not generalize to the complexity of dressed people in common images and videos. Additionally, current models lack the expres…

Cited by 435PDFcodeScholar
2019

Capture, Learning, and Synthesis of 3D Speaking Styles

CVPR 2019poster

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we introduce a unique 4D face dataset with about 29 minutes of 4D…

Cited by 426PDFcodeScholar
2019

Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation

CVPR 2019poster

We address the unsupervised learning of several interconnected problems in low-level vision: single view depth prediction, camera motion estimation, optical flow, and segmentation of a video into the static scene and moving regions. Our key insight is that these four fundamental vision problems are…

Cited by 742PDFcodeScholar
2018

Generating 3D Faces using Convolutional Mesh Autoencoders

ECCV 2018poster

Learned 3D representations of human faces are useful for computer vision problems such as 3D face tracking and reconstruction from images, as well as graphics applications such as character generation and animation. Traditional models learn a latent representation of a face using linear subspaces or…

2018

Unsupervised Learning of Multi-Frame Optical Flow with Occlusions

ECCV 2018poster

Learning optical flow with neural networks is hampered by the need for obtaining training data with associated ground truth. Unsupervised learning is a promising direction, yet the performance of current unsupervised methods is still limited. In particular, the lack of proper occlusion handling in c…

Cited by 232SourcePDFScholar