← Search

Yusu Qian

3 accepted papers

2026

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

CVPR 2026

Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community's progress remains constrained by the absence of large-scale, high-quality, and openly accessible

Cited by 0SourcecodeScholar
2025

MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

ICLR 2025poster

Effective evaluation of Multimodal Large Language Models (MLLMs) is essential for understanding their capabilities and limitations. In this paper, we introduce MIA-Bench, a benchmark designed to assess MLLMs’ ability to strictly adhere to complex instructions. Our benchmark comprises a diverse set o…

2025

UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

ICCV 2025poster

Text-to-Image (T2I) diffusion models have shown impressive results in generating visually compelling images following user prompts. Building on this, various methods further fine-tune the pre-trained T2I model for specific tasks. However, this requires separate model architectures, training designs,…