2026
VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA
CVPR 2026
Real-world documents combine text with tables, charts, photographs, and diagrams arranged in diverse layouts, yet existing research on multimodal large language models (MLLMs) for document QA predominantly produces text-only responses, underutilizing these visual elements. We introduce VinQA, a data