← Search

Zixiang Zhou

17 accepted papers

2026

ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars

CVPR 2026

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We pres

Cited by 0SourceScholar
2026

SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models

CVPR 2026

Post-training alignment of video generation models with human preferences is a critical goal. Developing effective Reward Models (RMs) for this process faces significant methodological hurdles. Current data collection paradigms, reliant on in-prompt pairwise annotations, suffer from labeling noise.

Cited by 0SourcecodeScholar
2026

StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars

CVPR 2026

Real-time, streaming interactive avatars represent a critical yet challenging goal in digital human research. Although diffusion-based human avatar generation methods achieve remarkable success, their non-causal architecture and high computational costs make them unsuitable for streaming. Moreover,

Cited by 0SourcecodeScholar
2026

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

CVPR 2026

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a unified framework for human-centric joint audio and video ge

Cited by 0SourceScholar
2025

Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

ICCV 2025poster

Talking head synthesis is vital for virtual avatars and human-computer interaction. However, most existing methods are typically limited to accepting control from a single primary modality, restricting their practical utility. To this end, we introduce ACTalker, an end-to-end video diffusion framewo…

2025

HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation

CVPR 2025poster

We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the chara…

2025

Inchworm-Inspired Bipedal Crawling Soft Robot With Forward and Backward Locomotion for Confined Spaces

RA-L 2025

Soft crawling robots have broad application prospects in rescue, exploration, and medical fields. Currently, many bionic soft crawling robots achieve multi-directional locomotion through steering mechanisms. However, their steering capabilities are limited when crawling in pipes or confined spaces,

Cited by 3SourceScholar
2024

AvatarGPT: All-in-One Framework for Motion Understanding Planning Generation and Beyond

CVPR 2024poster

Large Language Models(LLMs) have shown remarkable emergent abilities in unifying almost all (if not every) NLP tasks. In the human motion-related realm however researchers still develop siloed models for each task. Inspired by InstuctGPT[??] and the generalist concept behind Gato [??] we introduce A…

Cited by 32SourcePDFScholar
2024

Inchworm-Inspired Soft Robot With Controllable Locomotion Based on Self-Sensing of Deformation

RA-L 2024

Inchworm-inspired robots have become a prominent fixture in bionic research, mainly owing to the hotspot's focus on manufacturing actuating materials and bionic structures. An inchworm can crawl stably along a contact surface using gait control, achieved through muscle actuation and accurate percept

Cited by 14SourceScholar
2024

LPFormer: LiDAR Pose Estimation Transformer with Multi-Task Network

ICRA 2024poster

Due to the difficulty of acquiring large-scale 3D human keypoint annotation, previous methods for 3D human pose estimation (HPE) have often relied on 2D image features and sequential 2D annotations. Furthermore, the training of these networks typically assumes the prediction of a human bounding box…

Cited by 11SourceScholar
2024

LiDARFormer: A Unified Transformer-based Multi-task Network for LiDAR Perception

ICRA 2024poster

There is a recent need in the LiDAR perception field for unifying multiple tasks in a single strong network with improved performance, as opposed to using separate networks for each task. In this paper, we introduce a new LiDAR multi-task learning paradigm based on the transformer. The proposed LiDA…

Cited by 10SourceScholar
2023

LAMP: Leveraging Language Prompts for Multi-Person Pose Estimation

IROS 2023poster

Human-centric visual understanding is an important desideratum for effective human-robot interaction. In order to navigate crowded public places, social robots must be able to interpret the activity of the surrounding humans. This paper addresses one key aspect of human-centric visual understanding,…

Cited by 6SourcecodeScholar
2023

LidarMultiNet: Towards a Unified Multi-Task Network for LiDAR Perception

AAAI 2023technical

LiDAR-based 3D object detection, semantic segmentation, and panoptic segmentation are usually implemented in specialized networks with distinctive architectures that are difficult to adapt to each other. This paper presents LidarMultiNet, a LiDAR-based multi-task network that unifies these three maj…

Cited by 93SourcePDFScholar
2022

CenterFormer: Center-based Transformer for 3D Object Detection

ECCV 2022poster

"Query-based transformer has shown great potential in constructing long-range attention in many image-domain tasks, but has rarely been considered in LiDAR-based 3D object detection due to the overwhelming size of the point cloud data. In this paper, we propose CenterFormer, a center-based transform…

2021

Panoptic-PolarNet: Proposal-Free LiDAR Point Cloud Panoptic Segmentation

CVPR 2021poster

Panoptic segmentation presents a new challenge in exploiting the merits of both detection and segmentation, with the aim of unifying instance segmentation and semantic segmentation in a single framework. However, an efficient solution for panoptic segmentation in the emerging domain of LiDAR point c…

Cited by 149PDFcodeScholar
2020

PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation

CVPR 2020poster

The requirement of fine-grained perception by autonomous driving systems has resulted in recently increased research in the online semantic segmentation of single-scan LiDAR. Emerging datasets and technological advancements have enabled researchers to benchmark this problem and improve the applicabl…

Cited by 630PDFcodeScholar