← Search

Yalong Bai

15 accepted papers

2026

Masked Region Transformer for Layered Image Generation and Editing at Scale

CVPR 2026

Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editing in natural language. Despite its importance, this remains an underexplored area at scale. To address this gap, we pres

Cited by 0SourceScholar
2026

Pareto-Guided Optimal Transport for Multi-Reward Alignment

ICML 2026poster

Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward models remains a significant challenge. Existing multi-reward fusion approaches rely on weighted summation, which is costly to tune and insufficient for …

Cited by 0SourceScholar
2026

PreferThinker: Reasoning-based Personalized Image Preference Assessment

ICLR 2026poster

Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing methods mainly focus on general preference assessment, training models with large-scale data to tackle well-defined task…

Cited by 0SourceScholar
2026

Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation

CVPR 2026

Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on multimodal large language models to infer user preferences, but the derived prompts or latent codes rarely reflect them faithfully, leading to suboptim

Cited by 0SourceScholar
2025

Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment

ICCV 2025poster

Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have failed to evolve in parallel. This study reveals that human preference reward models fine-tuned based on CLIP and BLIP arch…

2024

CAMEL: CAusal Motion Enhancement Tailored for Lifting Text-driven Video Editing

CVPR 2024poster

Text-driven video editing poses significant challenges in exhibiting flicker-free visual continuity while preserving the inherent motion patterns of original videos. Existing methods operate under a paradigm where motion and appearance are intricately intertwined. This coupling leads to the network…

Cited by 4SourcePDFScholar
2024

Dynamic Prompt Optimizing for Text-to-Image Generation

CVPR 2024poster

Text-to-image generative models specifically those based on diffusion models like Imagen and Stable Diffusion have made substantial advancements. Recently there has been a surge of interest in the delicate refinement of text prompts. Users assign weights or alter the injection time steps of certain…

2022

Responsive Listening Head Generation: A Benchmark Dataset and Baseline

ECCV 2022poster

"We present a new listening head generation benchmark, for synthesizing responsive feedbacks of a listener (e.g., nod, smile) during a face-to-face conversation. As the indispensable complement to talking heads generation, listening head generation has seldomly been studied in literature. Automatica…

Cited by 60SourcePDFScholar
2021

Exploiting Relationship for Complex-scene Image Generation

AAAI 2021technical

The significant progress on Generative Adversarial Networks (GANs) has facilitated realistic single-object image generation based on language input. However, complex-scene generation (with various interactions among multiple objects) still suffers from messy layouts and object distortions, due to di…

2020

Look-Into-Object: Self-Supervised Structure Modeling for Object Recognition

CVPR 2020poster

Most object recognition approaches predominantly focus on learning discriminative visual patterns, while overlooking the holistic object structure. Though important, structure modeling usually requires significant manual annotations and therefore is labor-intensive. In this paper, we propose to "loo…

Cited by 99PDFcodeScholar
2019

Destruction and Construction Learning for Fine-Grained Image Recognition

CVPR 2019poster

Delicate feature representation about object parts plays a critical role in fine-grained recognition. For example, experts can even distinguish fine-grained objects relying only on object parts according to professional knowledge. In this paper, we propose a novel "Destruction and Construction Learn…

Cited by 599PDFcodeScholar
2018

Deep Attention Neural Tensor Network for Visual Question Answering

ECCV 2018poster

Visual question answering (VQA) has drawn great attention in cross-modal learning problems, which enables a machine to answer a natural language question given a reference image. Significant progress has been made by learning rich embedding features from images and questions by bilinear models, whil…

Cited by 83SourcePDFScholar