← Search

Jeonghyung Park

1 accepted papers

2024

Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding

NeurIPS 2024poster

We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effectively handle various font sizes within document images. To address the increasing costs of considering the multi-scale…

Cited by 2SourcePDFScholar