← Search

Amy Zhao

11 accepted papers

2026

BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

CVPR 2026

Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-language modeling that extensively improves upon BabyVLM-V1 through a longitudinal,

Cited by 0SourcecodeScholar
2026

EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR

CVPR 2026

Egocentric 3D human motion estimation is essential for AR/VR experiences, yet remains challenging due to limited body coverage from the egocentric viewpoint, frequent occlusions, and scarce labeled data. We present EgoPoseFormer v2, a method that addresses these challenges through two key contributi

Cited by 0SourceScholar
2026

Geometric Neural Distance Fields for Learning Human Motion Priors

CVPR 2026

We introduce Neural Riemannian Motion Fields (\name), a novel 3D generative human motion prior that enables robust, temporally consistent, and physically plausible 3D motion recovery. Unlike existing VAE or diffusion-based methods, our higher-order motion prior explicitly models the human motion in

Cited by 0SourceScholar
2026

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens

CVPR 2026

Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language generation and understanding remains largely underexplored. Existing approaches often fine-tune large language models (LLM

Cited by 0SourcecodeScholar
2024

EgoBody3M: Egocentric Body Tracking on a VR Headset using a Diverse Dataset

ECCV 2024poster

"Accurate tracking of a user’s body pose while wearing a virtual reality (VR), augmented reality (AR) or mixed reality (MR) headset is a prerequisite for authentic self-expression, natural social presence, and intuitive user interfaces. Existing body tracking approaches on VR/AR devices are either u…

Cited by 4SourcePDFScholar
2023

Social Diffusion: Long-term Multiple Human Motion Anticipation

ICCV 2023poster

We propose Social Diffusion, a novel method for short-term and long-term forecasting of the motion of multiple persons as well as their social interactions. Jointly forecasting motions for multiple persons involved in social activities is inherently a challenging problem due to the interdependenci…

Cited by 20PDFcodeScholar
2020

Painting Many Pasts: Synthesizing Time Lapse Videos of Paintings

CVPR 2020poster

We introduce a new video synthesis task: synthesizing time lapse videos depicting how a given painting might have been created. Artists paint using unique combinations of brushes, strokes, and colors. There are often many possible ways to create a given painting. Our goal is to learn to capture this…

Cited by 13PDFScholar
2019

Data Augmentation Using Learned Transformations for One-Shot Medical Image Segmentation

CVPR 2019oral

Image segmentation is an important task in many medical applications. Methods based on convolutional neural networks attain state-of-the-art accuracy; however, they typically rely on supervised training with large labeled datasets. Labeling medical images requires significant expertise and time, and…

Cited by 608PDFcodeScholar
2019

Visual Deprojection: Probabilistic Recovery of Collapsed Dimensions

ICCV 2019poster

We introduce visual deprojection: the task of recovering an image or video that has been collapsed along a dimension. Projections arise in various contexts, such as long-exposure photography, where a dynamic scene is collapsed in time to produce a motion-blurred image, and corner cameras, where refl…

Cited by 16PDFScholar
2018

An Unsupervised Learning Model for Deformable Medical Image Registration

CVPR 2018poster

We present a fast learning-based algorithm for deformable, pairwise 3D medical image registration. Current registration methods optimize an objective function independently for each pair of images, which can be time-consuming for large data. We define registration as a parametric function, and optim…

2018

Synthesizing Images of Humans in Unseen Poses

CVPR 2018poster

We address the computational problem of novel human pose synthesis. Given an image of a person and a desired pose, we produce a depiction of that person in that pose, retaining the appearance of both the person and background. We present a modular generative neural network that synthesizes unseen po…

Cited by 376SourcePDFScholar