← Search

Yuanzhi Liang

9 accepted papers

2026

Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation

CVPR 2026

Group Relative Policy Optimization (GRPO) has emerged as an effective and lightweight framework for post-training visual generative models. However, its performance is fundamentally limited by the ambiguity of textual-visual correspondence: a single prompt may validly describe diverse visual outputs

Cited by 0SourceScholar
2026

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

AAAI 2026technical

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, whi

Cited by 0SourcePDFScholar
2026

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation

CVPR 2026

Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, existing GRPO pipelines rely on a single scalar reward per sample, treating each im

Cited by 0SourceScholar
2024

FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention

NeurIPS 2024poster

Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing long video diffusion models. This paper investigates a strai…

Cited by 22SourcePDFScholar
2023

MAAL: Multimodality-Aware Autoencoder-Based Affordance Learning for 3D Articulated Objects

ICCV 2023poster

Inferring affordance for 3D articulated objects is a challenging and practical problem. It is a primary problem for applying robots to real-world scenarios. The exploration can be summarized as figuring out where to act and how to act. Correspondingly, the task mainly requires producing actionabilit…

Cited by 3PDFcodeScholar
2022

A Simple Episodic Linear Probe Improves Visual Recognition in the Wild

CVPR 2022poster

Understanding network generalization and feature discrimination is an open research problem in visual recognition. Many studies have been conducted to assess the quality of feature representations. One of the simple strategies is to utilize a linear probing classifier to quantitatively evaluate the…

Cited by 19PDFcodeScholar
2022

SEEG: Semantic Energized Co-Speech Gesture Generation

CVPR 2022poster

Talking gesture generation is a practical yet challenging task which aims to synthesize gestures in line with speech. Gestures with meaningful signs can better convey useful information and arouse sympathy in the audience. Current works focus on aligning gestures with the speech rhythms, which are h…

Cited by 58PDFcodeScholar