← Search

Xincheng Shuai

7 accepted papers

2026

Free-Form Scene Editor: Enabling Multi-Round Object Manipulation Like in a 3D Engine

AAAI 2026technical

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation. In this work, we present FFSE, a 3D-aware autoregressive framework designed to enable intuitive, physically-consistent o

Cited by 0SourcePDFScholar
2026

GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering

CVPR 2026

Generating accurate glyphs for visual text rendering is essential yet challenging. Existing methods typically enhance text rendering by training on a large amount of high-quality scene text images, but the limited coverage of glyph variations and excessive stylization often compromise glyph accuracy

Cited by 0SourcecodeScholar
2026

PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow

CVPR 2026

Graphic design is a creative and innovative process that plays a crucial role in applications such as e-commerce and advertising. However, developing an automated design system that can faithfully translate user intentions into editable design files remains an open challenge. Although recent studies

Cited by 0SourcecodeScholar
2026

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

ICML 2026poster

Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically assess understanding and generation capabilities in isolation, overlooking the synergy between comprehension and generation.…

Cited by 0SourceScholar
2025

Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation

ICCV 2025poster

Controlling the movements of dynamic objects and the camera within generated videos is a meaningful yet challenging task. Due to the lack of datasets with comprehensive 6D pose annotations, existing text-to-video methods can not simultaneously control the motions of both camera and objects in 3D-awa…

Cited by 0SourcePDFScholar
2025

SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation

NeurIPS 2025spotlight

Controllable image generation has attracted increasing attention in recent years, enabling users to manipulate visual content such as identity and style. However, achieving simultaneous control over the 9D poses (location, size, and orientation) of multiple objects remains an open challenge. Despite…

Cited by 0SourceScholar