← Search

Yiji Cheng

8 accepted papers

2026

Meta-CoT: Enhancing Granularity and Generalization in Image Editing

CVPR 2026

Unified multi-modal understanding/generative models have shown improved image editing performance by incorporating fine-grained understanding into their Chain-of-Thought (CoT) process. However, a critical question remains underexplored: what forms of CoT and training strategy can jointly enhance bot

Cited by 0SourcecodeScholar
2026

PromptEnhancer: Taming Your Rewriter for Text-to-Image Generation via Fine-Grained Reward

CVPR 2026

Recent text-to-image (T2I) diffusion models have achieved impressive progress in generating high-fidelity images, yet they often fail to faithfully follow complex user prompts, especially in attribute binding, negation, and compositional reasoning. To address this limitation, we propose PromptEnhanc

Cited by 0SourcecodeScholar
2026

Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing

CVPR 2026

In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models exhibit promising understanding capabilities, these strengt

Cited by 0SourceScholar
2026

TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts

CVPR 2026

Unified image generation and editing models suffer from severe task interference in dense diffusion transformers architectures, where a shared parameter space must compromise between conflicting objectives (e.g., local editing v.s. subject-driven generation). While the sparse Mixture-of-Experts (MoE

Cited by 0SourcecodeScholar
2025

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

CVPR 2025poster

Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of visual language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-gra…

Cited by 2SourcePDFScholar
2024

"Plan, Posture and Go: Towards Open-vocabulary Text-to-Motion Generation"

ECCV 2024poster

"Conventional text-to-motion generation methods are usually trained on limited text-motion pairs, making them hard to generalize to open-vocabulary scenarios. Some works use the CLIP model to align the motion space and the text space, aiming to enable motion generation from natural language motion d…

Cited by 1SourcePDFScholar
2024

GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling

NeurIPS 2024poster

We introduce a radiance representation that is both structured and fully explicit and thus greatly facilitates 3D generative modeling. Existing radiance representations either require an implicit feature decoder, which significantly degrades the modeling power of the representation, or are spatially…

Cited by 9SourcePDFScholar
2024

RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models

ECCV 2024poster

"We present RodinHD, which can generate high-fidelity 3D avatars from a portrait image. Existing methods fail to capture intricate details such as hairstyles which we tackle in this paper. We first identify an overlooked problem of catastrophic forgetting that arises when fitting triplanes sequentia…