← Search

Qirui Jiao

3 accepted papers

2026

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

ICML 2026poster

While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the long, detailed prompts required for professional applications. We present DetailMaster, a comprehensive benchmark for evaluating T2I capabilities on long prompt…

Cited by 0SourcecodeScholar
2026

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks

CVPR 2026

Evaluating the nuanced human-centric video understanding capabilities of Multimodal Large Language Models (MLLMs) remains a great challenge, as existing benchmarks often overlook the intricacies of emotion, behavior, and cross-modal alignment. We introduce HumanVBench, a comprehensive video benchmar

Cited by 0SourcecodeScholar
2025

Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

CVPR 2025poster

High-performance Multimodal Large Language Models (MLLMs) rely heavily on data quality. This study introduces a novel data synthesis method, leveraging insights from contrastive learning and image difference captioning to enhance fine-grained image recognition in MLLMs. By analyzing object differenc…

Cited by 11SourcePDFScholar