← Search

Sicheng Mo

7 accepted papers

2026

Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration

CVPR 2026

In this work, we explore an untapped signal in diffusion model inference. While all previous methods generate images independently at inference, we instead ask if samples can be generated collaboratively. We propose Group Diffusion, unlocking the attention mechanism to be shared across images, rathe

Cited by 0SourcecodeScholar
2025

X-Fusion: Introducing New Modality to Frozen Large Language Models

ICCV 2025poster

We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tower design with modality-specific weights, keeping the LLM's parameters frozen while integrating vision-specific informat…

Cited by 0SourcePDFScholar
2024

Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance

NeurIPS 2024poster

Recent controllable generation approaches such as FreeControl and Diffusion Self-Guidance bring fine-grained spatial and appearance control to text-to-image (T2I) diffusion models without training auxiliary modules. However, these methods optimize the latent embedding for each type of score function…

2024

FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition

CVPR 2024poster

Recent approaches such as ControlNet offer users fine-grained spatial control over text-to-image (T2I) diffusion models. However auxiliary modules have to be trained for each spatial condition type model architecture and checkpoint putting them at odds with the diverse intents and preferences a huma…

2024

SimGen: Simulator-conditioned Driving Scene Generation

NeurIPS 2024poster

Controllable synthetic data generation can substantially lower the annotation cost of training data. Prior works use diffusion models to generate driving images conditioned on the 3D object layout. However, those models are trained on small-scale datasets like nuScenes, which lack appearance and lay…

Cited by 9SourcePDFScholar