2024
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
ACL 2024findings
This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fine-grained and coarse-grained levels by facilitating a nuanced correlation betwe…