← Search

Kangrui Cen

3 accepted papers

2026

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation

ICML 2026poster

Unified Multimodal Models (UMMs) integrate both visual understanding and generation within a single framework. Their ultimate aspiration is to create a cycle where understanding and generation mutually reinforce each other. While recent post-training methods have successfully leveraged understanding…

Cited by 0SourceScholar
2026

LayerT2V: A Unified Multi-Layer Video Generation Framework

ICML 2026poster

Text-to-video generation has advanced rapidly, but existing methods typically output only the final composited video and lack editable layered representations, limiting their use in professional workflows. We propose \textbf{LayerT2V}, a unified multi-layer video generation framework that produces m…

Cited by 2SourceScholar
2026

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition

CVPR 2026

In controllable image generation, synthesizing coherent and consistent images from multiple reference inputs, i.e., **Multi-Image Composition** (MICo), remains a challenging problem, partly hindered by the lack of high-quality training data.To bridge this gap, we conduct a systematic study of MICo,

Cited by 0SourcecodeScholar