2026
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
CVPR 2026
Recently, multimodal large language models (MLLMs) have emerged as a unified paradigm for language and image generation. Compared with diffusion models, MLLMs possess a much stronger capability for semantic understanding, enabling them to process more complex textual inputs and comprehend richer con