← Search

Jinyang Tang

3 accepted papers

2026

AVGGT: Rethinking Global Attention for Accelerating VGGT

CVPR 2026

Models such as VGGT and \pi^3 have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a systematic analysis of how global attention contributes to multi-view

Cited by 0SourceScholar
2023

LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding

ACL 2023long

Visually-rich Document Understanding (VrDU) has attracted much research attention over the past years. Pre-trained models on a large number of document images with transformer-based backbones have led to significant performance gains in this field. The major challenge is how to fusion the different…

2023

Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction

EMNLP 2023long main

Recent advances in multimodal pre-trained models have significantly improved information extraction from visually-rich documents (VrDs), in which named entity recognition (NER) is treated as a sequence-labeling task of predicting the BIO entity tags for tokens, following the typical setting of NLP.…

Cited by 0SourcecodeScholar