2024
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
NeurIPS 2024poster
We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effectively handle various font sizes within document images. To address the increasing costs of considering the multi-scale…