← Search

Taichi Iki

3 accepted papers

2025

VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents

CVPR 2025poster

We aim to develop a retrieval-augmented generation (RAG) framework that answers questions over a corpus of visually-rich documents presented in mixed modalities (e.g., charts, tables) and diverse formats (e.g., PDF, PPTX). In this paper, we introduce a new RAG framework, VDocRAG, which can directly…

Cited by 3SourcePDFScholar
2024

InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions

AAAI 2024technical

We study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-written instructions. To this end, we propose InstructDoc, the first large-scale collection of 30 publicly available VDU da…

2021

Effect of Visual Extensions on Natural Language Understanding in Vision-and-Language Models

EMNLP 2021main

A method for creating a vision-and-language (V&L) model is to extend a language model through structural modifications and V&L pre-training. Such an extension aims to make a V&L model inherit the capability of natural language understanding (NLU) from the original language model. To see how well thi…