← Search

Zhenyuan Qin

3 accepted papers

2026

Free-Form Scene Editor: Enabling Multi-Round Object Manipulation Like in a 3D Engine

AAAI 2026technical

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation. In this work, we present FFSE, a 3D-aware autoregressive framework designed to enable intuitive, physically-consistent o

Cited by 0SourcePDFScholar
2025

Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation

ICCV 2025poster

Controlling the movements of dynamic objects and the camera within generated videos is a meaningful yet challenging task. Due to the lack of datasets with comprehensive 6D pose annotations, existing text-to-video methods can not simultaneously control the motions of both camera and objects in 3D-awa…

Cited by 0SourcePDFScholar
2025

SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation

NeurIPS 2025spotlight

Controllable image generation has attracted increasing attention in recent years, enabling users to manipulate visual content such as identity and style. However, achieving simultaneous control over the 9D poses (location, size, and orientation) of multiple objects remains an open challenge. Despite…

Cited by 0SourceScholar