← Search

Jidong Jiang

5 accepted papers

2026

Consis-GCPO: Consistency-Preserving Group Causal Preference Optimization for Vision Customization

ICLR 2026poster

Subject-driven generation faces a fundamental challenge: achieving high subject fidelity while maintaining semantic alignment with textual descriptions. While recent GRPO-based approaches have shown promise in aligning generative models with human preferences, they apply uniform optimization across…

Cited by 0SourceScholar
2026

FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus

AAAI 2026technical

Multi-subject personalized image generation aims to synthesize customized images containing multiple specified subjects without requiring test-time optimization. However, achieving fine-grained independent control over multiple subjects remains challenging due to difficulties in preserving subject f

Cited by 0SourcePDFScholar
2026

MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement

ICLR 2026poster

Multi-subject personalized generation presents unique challenges in maintaining identity fidelity and semantic coherence when synthesizing images conditioned on multiple reference subjects. Existing methods often suffer from identity blending and attribute leakage due to inadequate modeling of how d…

Cited by 0SourcecodeScholar
2026

MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation

ICLR 2026poster

Long video generation with Diffusion Transformers (DiTs) is bottlenecked by the quadratic scaling of full attention with sequence length. Since attention is highly redundant, outputs are dominated by a small subset of query–key pairs. Existing sparse methods rely on blockwise coarse estimation, whos…

Cited by 0SourcecodeScholar
2026

Prompt Yourself: Awakening Textual Semantics in 1D Visual Tokenizers

CVPR 2026

One-dimensional (1D) visual tokenizers offer notable semantic compactness by discarding local spatial priors, and have become increasingly popular for image reconstruction and generation tasks. However, such global and sequential representations struggle to preserve fine-grained visual content; simp

Cited by 0SourceScholar