← Search

Paolo Garza

2 accepted papers

2024

3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

ACL 2024findings

This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fine-grained and coarse-grained levels by facilitating a nuanced correlation betwe…

2024

Beyond Accuracy Optimization: Computer Vision Losses for Large Language Model Fine-Tuning

EMNLP 2024finding

Large Language Models (LLMs) have demonstrated impressive performance across various tasks. However, current training approaches combine standard cross-entropy loss with extensive data, human feedback, or ad hoc methods to enhance performance. These solutions are often not scalable or feasible due t…