← Search

Angjoo Kanazawa

73 accepted papers

2026

Flow Matching Policy Gradients

ICLR 2026poster

Flow-based generative models, including diffusion models, excel at modeling continuous distributions in high-dimensional spaces. In this work, we introduce Flow Policy Optimization (FPO), a simple on-policy reinforcement learning algorithm that brings flow matching into the policy gradient framework…

Cited by 0SourcecodeScholar
2026

OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction

ICRA 2026poster

A dominant paradigm for teaching humanoid robots complex skills is to retarget human motions as kinematic references to train reinforcement learning (RL) policies. However, existing retargeting pipelines often struggle with the significant embodiment gap between humans and robots, producing physical…

2026

Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching

RSS 2026poster

While recent advances in humanoid locomotion have achieved stable walking on varied terrains, capturing the agility and adaptivity of highly dynamic human motions remains an open challenge. In particular, agile parkour in complex environments demands not only low-level robustness, but also human-lik…

Cited by 0SourceScholar
2026

TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System

ICRA 2026poster

Large-scale data has driven breakthroughs in robotics, from language models to vision-language-action models in bimanual manipulation. However, humanoid robotics lacks equally effective data collection frameworks. Existing humanoid teleoperation systems either use decoupled control or depend on expe…

2026

Viser: Imperative, Web-based 3D Visualization for Python

RSS 2026poster

We present Viser, a toolkit for 3D visualization in robotics and computer vision. Viser aims to bring easy and extensible 3D visualization to Python: we provide comprehensive 3D scene and 2D GUI primitives, which can be used independently with minimal setup or composed to build specialized interface…

Cited by 0SourceScholar
2025

Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos

ICLR 2025poster

We present Agent-to-Sim (ATS), a framework for learning interactive behavior models of 3D agents from casual longitudinal video collections. Different from prior works that rely on marker-based tracking and multiview cameras, ATS learns natural behaviors of animal agents non-invasively through video…

2025

Continuous 3D Perception Model with Persistent State

CVPR 2025poster

We present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this evolving state can be used to generate metric-scale pointmaps (pe…

2025

Estimating Body and Hand Motion in an Ego-sensed World

CVPR 2025highlight

We present EgoAllo, a system for human motion estimation from a head-mounted device. Using only egocentric SLAM poses and images, EgoAllo guides sampling from a conditional diffusion model to estimate 3D body pose, height, and hand parameters that capture a device wearer's actions in the allocentric…

Cited by 5SourcePDFScholar
2025

Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop

CoRL 2025poster

Humans do not passively observe the visual world---we actively look in order to act. Motivated by this principle, we introduce EyeRobot, a robotic system with gaze behavior that emerges from the need to complete real-world tasks. We develop a mechanical eyeball that can freely rotate to observe its…

Cited by 0SourceScholar
2025

MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos

CVPR 2025award

We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input videos that feature predominantly static scenes with large…

Cited by 18SourcePDFScholar
2025

Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding

ICCV 2025poster

Whether snipping with scissors or opening a box, humans can quickly understand the 3D configurations of familiar objects. For novel objects, we can resort to long-form inspection to build intuition. The more we observe the object, the better we get at predicting its 3D state immediately. Existing sy…

Cited by 0SourcePDFScholar
2025

PyRoki: A Modular Toolkit for Robot Kinematic Optimization

IROS 2025

Robot motion can have many goals. Depending on the task, we might optimize for pose error, speed, collision, or similarity to a human demonstration. Motivated by this, we present PyRoki: a modular, extensible, and deviceagnostic toolkit for solving kinematic optimization problems. PyRoki couples an

Cited by 29SourcecodeScholar
2025

Reconstructing People, Places, and Cameras

CVPR 2025highlight

We present "Humans and Structure from Motion" (HSfM), a method for jointly reconstructing multiple human meshes, scene point clouds, and camera parameters in a metric world coordinate system from a sparse set of uncalibrated multi-view images featuring people. Our approach combines data-driven scene…

2025

Segment Any Motion in Videos

CVPR 2025poster

Moving object segmentation is a crucial task for achieving a high-level understanding of visual scenes and has numerous downstream applications. Humans can effortlessly segment moving objects in videos. Previous work has largely relied on optical flow to provide motion cues; however, this approach o…

2025

Shape of Motion: 4D Reconstruction from a Single Video

ICCV 2025poster

Monocular dynamic reconstruction is a challenging and long-standing vision problem due to the highly ill-posed nature of the task. Existing approaches depend on templates, are effective only in quasi-static scenes, or fail to model 3D motion explicitly. We introduce a method for reconstructing gener…

Cited by 0SourcePDFScholar
2025

St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World

ICCV 2025poster

Dynamic 3D reconstruction and point tracking in videos are typically treated as separate tasks, despite their deep connection. We propose St4RTrack, a feed-forward frame- work that simultaneously reconstructs and tracks dynamic video content in a world coordinate frame from RGB in- puts. This is ach…

Cited by 0SourcePDFScholar
2025

Visual Imitation Enables Contextual Humanoid Control

CoRL 2025oral

How can we teach humanoids to climb staircases and sit on chairs using the surrounding environment context? Arguably the simplest way is to _just show them_—casually capture a human motion video and feed it to humanoids. We introduce **VideoMimic**, a real-to-sim-to-real pipeline that mines everyday…

Cited by 0SourceScholar
2024

From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations

CVPR 2024poster

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio we output multiple possibilities of gestural motion for an individual including face body and hands. The key behind our method is…

2024

GARField: Group Anything with Radiance Fields

CVPR 2024poster

Grouping is inherently ambiguous due to the multiple levels of granularity in which one can decompose a scene --- should the wheels of an excavator be considered separate or part of the whole? We propose Group Anything with Radiance Fields (GARField) an approach for decomposing 3D scenes into a hier…

2024

Generative Proxemics: A Prior for 3D Social Interaction from Images

CVPR 2024poster

Social interaction is a fundamental aspect of human behavior and communication. The way individuals position themselves in relation to others also known as proxemics conveys social cues and affects the dynamics of social interaction. Reconstructing such interaction from images presents challenges be…

2024

NeRFiller: Completing Scenes via Generative 3D Inpainting

CVPR 2024poster

We propose NeRFiller an approach that completes missing portions of a 3D capture via generative 3D inpainting using off-the-shelf 2D visual generative models. Often parts of a captured 3D scene or object are missing due to mesh reconstruction failures or a lack of observations (e.g. contact regions…

Cited by 33SourcePDFScholar
2024

Reconstructing Hands in 3D with Transformers

CVPR 2024poster

We present an approach that can reconstruct hands in 3D from monocular input. Our approach for Hand Mesh Recovery HaMeR follows a fully transformer-based architecture and can analyze hands with significantly increased accuracy and robustness compared to previous work. The key to HaMeR's success lies…

2024

Rethinking Score Distillation as a Bridge Between Image Distributions

NeurIPS 2024poster

Score distillation sampling (SDS) has proven to be an important tool, enabling the use of large-scale diffusion priors for tasks operating in data-poor domains. Unfortunately, SDS has a number of characteristic artifacts that limit its utility in general-purpose applications. In this paper, we make…

Cited by 12SourcePDFScholar
2024

Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction

CoRL 2024poster

Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying new behaviors. This work develops Robot See Robot Do (RSRD), a method for imitating articulated object manipulation fr…

Cited by 15SourcecodeScholar
2024

The More You See in 2D the More You Perceive in 3D

CVPR 2024highlight

Humans can infer 3D structure from 2D images of an object based on past experience and improve their 3D understanding as they see more images. Inspired by this behavior we introduce SAP3D a system for 3D reconstruction and novel view synthesis from an arbitrary number of unposed images. Given a few…

2023

Can Language Models Learn to Listen?

ICCV 2023poster

We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the speaker's words with their timestamps, our approach autoregressively predicts a response of a listener: a sequence of lis…

Cited by 24PDFScholar
2023

Decoupling Human and Camera Motion From Videos in the Wild

CVPR 2023poster

We propose a method to reconstruct global human trajectories from videos in the wild. Our optimization method decouples the camera and human motion, which allows us to place people in the same world coordinate frame. Most existing methods do not model the camera motion; methods that rely on the back…

2023

Differentiable Blocks World: Qualitative 3D Decomposition by Rendering Primitives

NeurIPS 2023poster

Given a set of calibrated images of a scene, we present an approach that produces a simple, compact, and actionable 3D world representation by means of 3D primitives. While many approaches focus on recovering high-fidelity 3D scenes, we focus on parsing a scene into mid-level 3D representations made…

Cited by 18SourcePDFScholar
2023

Humans in 4D: Reconstructing and Tracking Humans with Transformers

ICCV 2023poster

We present an approach to reconstruct humans and track them over time. At the core of our approach, we propose a fully "transformerized" version of a network for human mesh recovery. This network, HMR 2.0, advances the state of the art and shows the capability to analyze unusual poses that have in t…

Cited by 221PDFcodeScholar
2023

Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions

ICCV 2023oral

We propose a method for editing NeRF scenes with text-instructions. Given a NeRF of a scene and the collection of images used to reconstruct it, our method uses an image-conditioned diffusion model (InstructPix2Pix) to iteratively edit the input images while optimizing the underlying scene, resultin…

Cited by 429PDFScholar
2023

K-Planes: Explicit Radiance Fields in Space, Time, and Appearance

CVPR 2023poster

We introduce k-planes, a white-box model for radiance fields in arbitrary dimensions. Our model uses d-choose-2 planes to represent a d-dimensional scene, providing a seamless way to go from static (d=3) to dynamic (d=4) scenes. This planar factorization makes adding dimension-specific priors easy,…

2023

Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping

CoRL 2023oral

Grasping objects by a specific subpart is often crucial for safety and for executing downstream tasks. We propose LERF-TOGO, Language Embedded Radiance Fields for Task-Oriented Grasping of Objects, which uses vision-language models zero-shot to output a grasp distribution over an object given a natu…

Cited by 88SourcecodeScholar
2023

Nerfbusters: Removing Ghostly Artifacts from Casually Captured NeRFs

ICCV 2023poster

Casually captured Neural Radiance Fields (NeRFs) suffer from artifacts such as floaters or flawed geometry when rendered outside the input camera trajectory. Existing evaluation protocols often do not capture these effects, since they usually only assess image quality at every 8th frame of the train…

Cited by 61PDFcodeScholar
2023

On the Benefits of 3D Pose and Tracking for Human Action Recognition

CVPR 2023poster

In this work we study the benefits of using tracking and 3D poses for action recognition. To achieve this, we take the Lagrangian view on analysing actions over a trajectory of human motion rather than at a fixed point in space. Taking this stand allows us to use the tracklets of people to predict t…

2022

Deformable Sprites for Unsupervised Video Decomposition

CVPR 2022oral

We describe a method to extract persistent elements of a dynamic scene from an input video. We represent each scene element as a Deformable Sprite consisting of three components: 1) a 2D texture image for the entire video, 2) per-frame masks for the element, and 3) non-rigid deformations that map th…

Cited by 78PDFScholar
2022

Differentiable Gradient Sampling for Learning Implicit 3D Scene Reconstructions from a Single Image

ICLR 2022poster

Implicit shape models are promising 3D representations for modeling arbitrary locations, with Signed Distance Functions (SDFs) particularly suitable for clear mesh surface reconstruction. Existing approaches for single object reconstruction impose supervision signals based on the loss of the signed…

Cited by 4SourcePDFScholar
2022

Evo-NeRF: Evolving NeRF for Sequential Robot Grasping of Transparent Objects

CoRL 2022oral

Sequential robot grasping of transparent objects, where a robot removes objects one by one from a workspace, is important in many industrial and household scenarios. We propose Evolving NeRF (Evo-NeRF), leveraging recent speedups in NeRF training and further extending it to rapidly train the NeRF re…

Cited by 100SourceScholar
2022

InfiniteNature-Zero: Learning Perpetual View Generation of Natural Scenes from Single Images

ECCV 2022poster

"We present a method for learning to generate unbounded flythrough videos of natural scenes starting from a single view. This capability is learned from a collection of single photographs, without requiring camera poses or even multiple views of each scene. To achieve this, we propose a novel self-s…

2022

Learning To Listen: Modeling Non-Deterministic Dyadic Facial Motion

CVPR 2022poster

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and speech audio of the speaker using a motion-audio cross attentio…

Cited by 105PDFScholar
2022

Monocular Dynamic View Synthesis: A Reality Check

NeurIPS 2022accept

We study the recent progress on dynamic view synthesis (DVS) from monocular video. Though existing approaches have demonstrated impressive results, we show a discrepancy between the practical capture process and the existing experimental protocols, which effectively leaks in multi-view signals durin…

2022

Plenoxels: Radiance Fields Without Neural Networks

CVPR 2022oral

We introduce Plenoxels (plenoptic voxels), a system for photorealistic view synthesis. Plenoxels represent a scene as a sparse 3D grid with spherical harmonics. This representation can be optimized from calibrated images via gradient methods and regularization without any neural components. On stand…

Cited by 1546PDFcodeScholar
2022

Studying Bias in GANs through the Lens of Race

ECCV 2022poster

"In this work, we study how the performance and evaluation of generative image models are impacted by the racial composition of the datasets upon which these models are trained. By examining and controlling the racial distributions in various training datasets, we are able to observe the impacts of…

Cited by 50SourcePDFScholar
2022

TAVA: Template-Free Animatable Volumetric Actors

ECCV 2022poster

"Coordinate-based volumetric representations have the potential to generate photo-realistic virtual avatars from images. However, virtual avatars need to be controllable and be rendered in novel poses that may not have been observed. Traditional techniques, such as LBS, provide such a controlling fu…

2022

The One Where They Reconstructed 3D Humans and Environments in TV Shows

ECCV 2022poster

"TV shows depict a wide variety of human behaviors and have been studied extensively for their potential to be a rich source of data for many applications. However, the majority of the existing work focuses on 2D recognition tasks. In this paper, we make the observation that there is a certain persi…

Cited by 28SourcePDFScholar
2022

Tracking People by Predicting 3D Appearance, Location and Pose

CVPR 2022oral

We present an approach for tracking people in monocular videos by predicting their future 3D representations. To achieve this, we first lift people to 3D from a single frame in a robust manner. This lifting includes information about the 3D pose of the person, their location in the 3D space, and the…

Cited by 74PDFcodeScholar
2021

AI Choreographer: Music Conditioned 3D Dance Generation With AIST++

ICCV 2021poster

We present AIST++, a new multi-modal dataset of 3D dance motion and music, along with FACT, a Full-Attention Cross-modal Transformer network for generating 3D dance motion conditioned on music. The proposed AIST++ dataset contains 1.1M frames of 3D dance motion in 1408 sequences, covering 10 dance g…

Cited by 576PDFcodeScholar
2021

De-Rendering the World's Revolutionary Artefacts

CVPR 2021poster

Recent works have shown exciting results in unsupervised image de-rendering--learning to decompose 3D shape, appearance, and lighting from single-image collections without explicit supervision. However, many of these assume simplistic material and lighting models. We propose a method, termed RADAR,…

Cited by 35PDFcodeScholar
2021

Infinite Nature: Perpetual View Generation of Natural Scenes From a Single Image

ICCV 2021poster

We introduce the problem of perpetual view generation - long-range generation of novel views corresponding to an arbitrarily long camera trajectory given a single image. This is a challenging problem that goes far beyond the capabilities of current view synthesis methods, which quickly degenerate wh…

Cited by 169PDFcodeScholar
2021

KeypointDeformer: Unsupervised 3D Keypoint Discovery for Shape Control

CVPR 2021poster

We introduce KeypointDeformer, a novel unsupervised method for shape control through automatically discovered 3D keypoints. We cast this as the problem of aligning a source 3D object to a target 3D object from the same object category. Our method analyzes the difference between the shapes of the two…

Cited by 70PDFcodeScholar
2021

PlenOctrees for Real-Time Rendering of Neural Radiance Fields

ICCV 2021poster

We introduce a method to render Neural Radiance Fields (NeRFs) in real time using PlenOctrees, an octree-based 3D representation which supports view-dependent effects. Our method can render 800x800 images at more than 150 FPS, which is over 3000 times faster than conventional NeRFs. We do so without…

Cited by 1173PDFcodeScholar
2021

Tracking People with 3D Representations

NeurIPS 2021poster

We present a novel approach for tracking multiple people in video. Unlike past approaches which employ 2D representations, we focus on using 3D representations of people, located in three-dimensional space. To this end, we develop a method, Human Mesh and Appearance Recovery (HMAR) which in addition…

2020

An Analysis of SVD for Deep Rotation Estimation

NeurIPS 2020poster

Symmetric orthogonalization via SVD, and closely related procedures, are well-known techniques for projecting matrices onto O(n) or SO(n). These tools have long been used for applications in computer vision, for example optimal 3D alignment problems solved by orthogonal Procrustes, rotation averagin…

2020

Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild

ECCV 2020poster

We present a method that infers spatial arrangements and shapes of humans and objects in a globally consistent 3D scene, all from a single image in-the-wild captured in an uncontrolled environment. Notably, our method runs on datasets without any scene- or object-level 3D supervision. Our key insigh…

2019

PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization

ICCV 2019poster

We introduce Pixel-aligned Implicit Function (PIFu), an implicit representation that locally aligns pixels of 2D images with the global context of their corresponding 3D object. Using PIFu, we propose an end-to-end deep learning method for digitizing highly detailed clothed humans that can infer bot…

Cited by 1438PDFScholar
2019

Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"

ICCV 2019poster

We present the first method to perform automatic 3D pose, shape and texture capture of animals from images acquired in-the-wild. In particular, we focus on the problem of capturing 3D information about Grevy's zebras from a collection of images. The Grevy's zebra is one of the most endangered specie…

Cited by 187PDFcodeScholar
2019

Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow

ICLR 2019poster

Adversarial learning methods have been proposed for a wide range of applications, but the training of adversarial models can be notoriously unstable. Effectively balancing the performance of the generator and discriminator is critical, since a discriminator that achieves very high accuracy will prod…

Cited by 282SourcePDFScholar
2018

End-to-End Recovery of Human Shape and Pose

CVPR 2018poster

We describe Human Mesh Recovery (HMR), an end-to-end framework for reconstructing a full 3D mesh of a human body from a single RGB image. In contrast to most current methods that compute 2D or 3D joint locations, we produce a richer and more useful mesh representation that is parameterized by shape…

2018

Learning Category-Specific Mesh Reconstruction from Image Collections

ECCV 2018poster

We present a learning framework for recovering the 3D shape, camera, and texture of an object from a single image. The shape is represented as a deformable 3D mesh model of an object category where a shape is parameterized by a learned mean shape and per-instance predicted deformation. Our approach…

2018

Lions and Tigers and Bears: Capturing Non-Rigid, 3D, Articulated Shape From Images

CVPR 2018poster

Animals are widespread in nature and the analysis of their shape and motion is important in many fields and industries. Modeling 3D animal shape, however, is difficult because the 3D scanning methods used to capture human shape are not applicable to wild animals or natural settings. Consequently, we…

Cited by 154SourcePDFScholar
2018

SfSNet: Learning Shape, Reflectance and Illuminance of Faces `in the Wild'

CVPR 2018poster

We present SfSNet, an end-to-end learning framework for producing an accurate decomposition of an unconstrained human face image into shape, reflectance and illuminance. SfSNet is designed to reflect a physical lambertian rendering model. SfSNet learns from a mixture of labeled synthetic and unlabel…

Cited by 377SourcePDFScholar
2017

3D Menagerie: Modeling the 3D Shape and Pose of Animals

CVPR 2017spotlight

There has been significant work on learning realistic, articulated, 3D models of the human body. In contrast, there are few such models of animals, despite many applications. The main challenge is that animals are much less cooperative than humans. The best human body models are learned from thousan…

Cited by 491PDFScholar