← Search

Yuanhao Zhai

10 accepted papers

2025

PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions

ICCV 2025poster

Diffusion-based generative models have shown promise in synthesizing histopathology images to address data scarcity caused by privacy constraints. Diagnostic text reports provide high-level semantic descriptions, and masks offer fine-grained spatial structures essential for representing distinct mor…

2025

SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation

ICLR 2025spotlight

Human beings are endowed with a complementary learning system, which bridges the slow learning of general world dynamics with fast storage of episodic memory from a new experience. Previous video generation models, however, primarily focus on slow learning by pre-training on vast amounts of data, ov…

Cited by 4SourcePDFScholar
2025

Text2Outfit: Controllable Outfit Generation with Multimodal Language Models

ICCV 2025poster

Existing outfit recommendation frameworks focus on outfit compatibility prediction and complementary item retrieval. We present a text-driven outfit generation framework, Text2Outfit, which generates outfits controlled by text prompts. Our framework supports two forms of outfit recommendation: 1) Te…

Cited by 0SourcePDFScholar
2024

DisCo: Disentangled Control for Realistic Human Dance Generation

CVPR 2024poster

Generative AI has made significant strides in computer vision particularly in text-driven image/video synthesis (T2I/T2V). Despite the notable advancements it remains challenging in human-centric content synthesis such as realistic dance generation. Current methodologies primarily tailored for human…

2024

Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation

NeurIPS 2024poster

Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, directly applying these techniques to video models results in unsatisfied frame quality. This issue arises from the limited frame appearance quality in public video datasets, affecting the performan…

2024

Seesaw: Compensating for Nonlinear Reduction with Linear Computations for Private Inference

ICML 2024poster

With increasingly serious data privacy concerns and strict regulations, privacy-preserving machine learning (PPML) has emerged to securely execute machine learning tasks without violating privacy. Unfortunately, the computational cost to securely execute nonlinear computations in PPML remains signif…

Cited by 5SourcePDFScholar
2023

High Fidelity 3D Hand Shape Reconstruction via Scalable Graph Frequency Decomposition

CVPR 2023poster

Despite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high fidelity hand modeling is required, e.g., personalized hand modeling. T…

2023

SOAR: Scene-debiasing Open-set Action Recognition

ICCV 2023poster

Deep models have the risk of utilizing spurious clues to make predictions, e.g., recognizing actions via classifying the background scene. This problem severely degrades the open-set action recognition performance when the testing samples exhibit scene distributions different from the training sampl…

Cited by 18PDFcodeScholar
2023

Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency Learning

ICCV 2023poster

As advanced image manipulation techniques emerge, detecting the manipulation becomes increasingly important. Despite the success of recent learning-based approaches for image manipulation detection, they typically require expensive pixel-level annotations to train, while exhibiting degraded performa…

Cited by 29PDFcodeScholar
2020

Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization

ECCV 2020poster

Weakly-supervised Temporal Action Localization (W-TAL) aims to classify and localize all action instances in an untrimmed video under only video-level supervision. However, without frame-level annotations, it is challenging for W-TAL methods to identify false positive action proposals and generate a…