← Search

Kyunghwan An

2 accepted papers

2026

VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA

CVPR 2026

Real-world documents combine text with tables, charts, photographs, and diagrams arranged in diverse layouts, yet existing research on multimodal large language models (MLLMs) for document QA predominantly produces text-only responses, underutilizing these visual elements. We introduce VinQA, a data

Cited by 0SourceScholar
2024

Structure-Aware Multimodal Sequential Learning for Visual Dialog

AAAI 2024technical

With the ability to collect vast amounts of image and natural language data from the web, there has been a remarkable advancement in Large-scale Language Models (LLMs). This progress has led to the emergence of chatbots and dialogue systems capable of fluent conversations with humans. As the variety…

Cited by 1SourcePDFScholar