← Search

Umar Iqbal

28 accepted papers

2026

EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

CVPR 2026

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion model trained on egocentric data. We train a video predicti

Cited by 0SourcecodeScholar
2025

AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview Diffusion

ICCV 2025poster

Existing methods for image-to-3D avatar generation struggle to produce highly detailed, animation-ready avatars suitable for real-world applications. We introduce AdaHuman, a novel framework that generates high-fidelity animatable 3D avatars from a single in-the-wild image. AdaHuman incorporates two…

Cited by 0SourcePDFScholar
2025

GENMO: A GENeralist Model for Human MOtion

ICCV 2025poster

Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes, while motion estimation models aim to reconstruct accurate mot…

Cited by 0SourcePDFScholar
2025

GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion

ICCV 2025poster

Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture fine-grained dynamic details. To address these limitations,…

Cited by 0SourcePDFScholar
2025

HumanOLAT: A Large-Scale Dataset for Full-Body Human Relighting and Novel-View Synthesis

ICCV 2025poster

Simultaneous relighting and novel-view rendering of digital human representations is an important yet challenging task with numerous applications. However, progress in this area has been significantly limited due to the lack of publicly available, high-quality datasets, especially for full-body huma…

Cited by 0SourcePDFScholar
2025

SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing

CVPR 2025poster

We introduce SimAvatar, a framework designed to generate simulation-ready clothed 3D human avatars from a text prompt. Current text-driven human avatar generation methods either model hair, clothing and human body using a unified geometry or produce hair and garments that are not easily adaptable fo…

Cited by 1SourcePDFScholar
2024

GAvatar: Animatable 3D Gaussian Avatars with Implicit Mesh Learning

CVPR 2024highlight

Gaussian splatting has emerged as a powerful 3D representation that harnesses the advantages of both explicit (mesh) and implicit (NeRF) 3D representations. In this paper we seek to leverage Gaussian splatting to generate realistic animatable avatars from textual descriptions addressing the limitati…

Cited by 41SourcePDFScholar
2024

What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs

CVPR 2024poster

3D-aware Generative Adversarial Networks (GANs) have shown remarkable progress in learning to generate multi-view-consistent images and 3D geometries of scenes from collections of 2D images via neural volume rendering. Yet the significant memory and computational costs of dense sampling in volume re…

Cited by 8SourcePDFScholar
2023

Generalizable One-shot 3D Neural Head Avatar

NeurIPS 2023poster

We present a method that reconstructs and animates a 3D head avatar from a single-view portrait image. Existing methods either involve time-consuming optimization for a specific person with multiple images, or they struggle to synthesize intricate appearance details beyond the facial region. To addr…

Cited by 31SourcePDFScholar
2023

Learning Human Dynamics in Autonomous Driving Scenarios

ICCV 2023poster

Simulation has emerged as an indispensable tool for scaling and accelerating the development of self-driving systems. A critical aspect of this is simulating realistic and diverse human behavior and intent. In this work, we propose a holistic framework for learning physically plausible human dynamic…

Cited by 23PDFScholar
2023

RANA: Relightable Articulated Neural Avatars

ICCV 2023poster

We propose RANA, a relightable and articulated neural avatar for the photorealistic synthesis of humans under arbitrary viewpoints, body poses, and lighting. We only require a short video clip of the person to create the avatar and assume no knowledge about the lighting environment. We present a nov…

Cited by 16PDFScholar
2022

GLAMR: Global Occlusion-Aware Human Mesh Recovery With Dynamic Cameras

CVPR 2022oral

We present an approach for 3D global human mesh recovery from monocular videos recorded with dynamic cameras. Our approach is robust to severe and long-term occlusions and tracks human bodies even when they go outside the camera's field of view. To achieve this, we first propose a deep generative mo…

Cited by 136PDFcodeScholar
2022

Watch It Move: Unsupervised Discovery of 3D Joints for Re-Posing of Articulated Objects

CVPR 2022poster

Rendering articulated objects while controlling their poses is critical to applications such as virtual reality or animation for movies. Manipulating the pose of an object, however, requires the understanding of its underlying structure, that is, its joints and how they interact with each other. Unf…

Cited by 51PDFcodeScholar
2021

DexYCB: A Benchmark for Capturing Hand Grasping of Objects

CVPR 2021poster

We introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estima…

Cited by 314PDFcodeScholar
2021

Learning to Track Instances without Video Annotations

CVPR 2021poster

Tracking segmentation masks of multiple instances has been intensively studied, but still faces two fundamental challenges: 1) the requirement of large-scale, frame-wise annotation, and 2) the complexity of two-stage approaches. To resolve these challenges, we introduce a novel semi-supervised frame…

Cited by 32PDFScholar
2021

Physics-Based Human Motion Estimation and Synthesis From Videos

ICCV 2021poster

Human motion synthesis is an important problem for applications in graphics and gaming, and even in simulation environments for robotics. Existing methods require accurate motion capture data for training, which is costly to obtain. Instead, we propose a framework for training generative models of p…

Cited by 110PDFScholar
2021

Self-Supervised Object Detection via Generative Image Synthesis

ICCV 2021poster

We present SSOD -- the first end-to-end analysis-by-synthesis framework with controllable GANs for the task of self-supervised object detection. We use collections of real-world images without bounding box annotations to learn to synthesize and detect objects. We leverage controllable GANs to synthe…

Cited by 15PDFcodeScholar
2021

Weakly-Supervised Physically Unconstrained Gaze Estimation

CVPR 2021poster

A major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environments are abundantly available and can be much more easily annotated with frame-…

Cited by 45PDFcodeScholar
2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

Self-Supervised Viewpoint Learning From Image Collections

CVPR 2020poster

Training deep neural networks to estimate the viewpoint of objects requires large labeled training datasets. However, manually labeling viewpoints is notoriously hard, error-prone, and time-consuming. On the other hand, it is relatively easy to mine many unlabeled images of an object category from t…

Cited by 45PDFcodeScholar
2020

Weakly Supervised 3D Hand Pose Estimation via Biomechanical Constraints

ECCV 2020poster

Estimating 3D hand pose from 2D images is a difficult, inverse problem due to the inherent scale and depth ambiguities. Current state-of-the-art methods train fully supervised deep neural networks with 3D ground-truth data. However, acquiring 3D annotations is expensive, typically requiring calibrat…

Cited by 185SourcePDFScholar
2019

Few-Shot Adaptive Gaze Estimation

ICCV 2019oral

Inter-personal anatomical differences limit the accuracy of person-independent gaze estimation networks. Yet there is a need to lower gaze errors further to enable applications requiring higher quality. Further gains can be achieved by personalizing gaze networks, ideally with few calibration sample…

Cited by 248PDFcodeScholar
2018

Hand Pose Estimation via Latent 2.5D Heatmap Regression

ECCV 2018poster

Estimating the 3D pose of a hand is an essential part of human-computer interaction. Estimating 3D pose using depth or multi-view sensors has become easier with recent advances in computer vision, however, regressing pose from a single RGB image is much less straightforward. The main difficulty aris…

Cited by 393SourcePDFScholar
2018

PoseTrack: A Benchmark for Human Pose Estimation and Tracking

CVPR 2018poster

Existing systems for video-based pose estimation and tracking struggle to perform well on realistic videos with multiple people and often fail to output body-pose trajectories consistent over time. To address this shortcoming this paper introduces PoseTrack which is a new large-scale benchmark for v…

Cited by 621SourcePDFScholar
2016

A Dual-Source Approach for 3D Pose Estimation From a Single Image

CVPR 2016spotlight

One major challenge for 3D pose estimation from a single RGB image is the acquisition of sufficient training data. In particular, collecting large amounts of training data that contain unconstrained images and are annotated with accurate 3D poses is infeasible. We therefore propose to use two indepe…

Cited by 260PDFScholar