← Search

Yizhong Zhang

8 accepted papers

2026

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

CVPR 2026

Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a framework for streamable talking portrait video generation conditioned on speech audio and reference images. Designed m

Cited by 0SourceScholar
2026

Scalable Vision-Language-Action Model Pretraining for Robotic Dexterous Manipulation with Real-Life Human Activity Videos

ICRA 2026poster

This paper presents an approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot end-effector, we show that "in-the-wild" egocentric human videos wit…

Cited by 0Scholar
2025

APA-BI: Adaptive Partition Aggregation and Bidirectional Integration for UAV-View Geo-Localization

ICRA 2025

The task of UAV-view geo-localization is to match a query image with database images to estimate the current geographic location of the query image. This is particularly useful in environments where GPS is not available or when the device fails. Although deep learning methods make sufficient progres

Cited by 1SourceScholar
2025

JRN-Geo: A Joint Perception Network Based on RGB and Normal Images for Cross-View Geo-Localization

ICRA 2025

Cross-view geo-localization plays a critical role in Unmanned Aerial Vehicle (UAV) localization and navigation. However, significant challenges arise from the drastic viewpoint differences and appearance variations between images. Existing methods predominantly rely on semantic features from RGB ima

Cited by 1SourceScholar
2025

Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning

ICRA 2025

Robotic manipulation in cluttered environments presents significant challenges, particularly when the clutter includes thin, deformable objects like cables, which complicate perception and decision-making processes. In the context of datacenters, the automation of networking tasks often involves the

Cited by 0SourceScholar
2025

UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping

CVPR 2025poster

We introduce UniGraspTransformer, a universal Transformer-based network for dexterous robotic grasping that simplifies training while enhancing scalability and performance. Unlike prior methods such as UniDexGrasp++, which require complex, multi-step training pipelines, UniGraspTransformer follows a…

2025

VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image

NeurIPS 2025poster

We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human faces, and reconstructing an intricate 3D head avatar from a single portrait image. To accurately model expression deta…

Cited by 0SourceScholar
2024

VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

NeurIPS 2024oral

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but als…

Cited by 92SourcePDFScholar