← Search

Mohan Zhou

5 accepted papers

2026

Masked Region Transformer for Layered Image Generation and Editing at Scale

CVPR 2026

Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editing in natural language. Despite its importance, this remains an underexplored area at scale. To address this gap, we pres

Cited by 0SourceScholar
2026

Pareto-Guided Optimal Transport for Multi-Reward Alignment

ICML 2026poster

Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward models remains a significant challenge. Existing multi-reward fusion approaches rely on weighted summation, which is costly to tune and insufficient for …

Cited by 0SourceScholar
2022

Responsive Listening Head Generation: A Benchmark Dataset and Baseline

ECCV 2022poster

"We present a new listening head generation benchmark, for synthesizing responsive feedbacks of a listener (e.g., nod, smile) during a face-to-face conversation. As the indispensable complement to talking heads generation, listening head generation has seldomly been studied in literature. Automatica…

Cited by 60SourcePDFScholar
2020

Look-Into-Object: Self-Supervised Structure Modeling for Object Recognition

CVPR 2020poster

Most object recognition approaches predominantly focus on learning discriminative visual patterns, while overlooking the holistic object structure. Though important, structure modeling usually requires significant manual annotations and therefore is labor-intensive. In this paper, we propose to "loo…

Cited by 99PDFcodeScholar