← Search

Young Rok Jang

2 accepted papers

2026

VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA

CVPR 2026

Real-world documents combine text with tables, charts, photographs, and diagrams arranged in diverse layouts, yet existing research on multimodal large language models (MLLMs) for document QA predominantly produces text-only responses, underutilizing these visual elements. We introduce VinQA, a data

Cited by 0SourceScholar
2025

Learning to Explore and Select for Coverage-Conditioned Retrieval-Augmented Generation

NAACL 2025findings

Interactions with large language models (LLMs) often yield long and detailed responses, leveraging both parametric knowledge and retrieval-augmented generation (RAG). While these responses can provide rich insights, they often include redundant or less engaging content not aligned with user interest…