← Search

Shangbang Long

4 accepted papers

2023

FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

ACL 2023long

The recent advent of self-supervised pre-training techniques has led to a surge in the use of multimodal learning in form document understanding. However, existing approaches that extend the mask language modeling to other modalities require careful multi-task tuning, complex reconstruction target d…

2022

Towards End-to-End Unified Scene Text Detection and Layout Analysis

CVPR 2022poster

Scene text detection and document layout analysis have long been treated as two separate tasks in different image domains. In this paper, we bring them together and introduce the task of unified scene text detection and layout analysis. The first hierarchical scene text dataset is introduced to enab…

Cited by 108PDFcodeScholar
2020

A New Perspective for Flexible Feature Gathering in Scene Text Recognition Via Character Anchor Pooling

ICASSP 2020accepted

Irregular scene text recognition has attracted much attention from the research community, mainly due to the complexity of shapes of text in natural scene. However, recent methods either rely on shape-sensitive modules such as bounding box regression, or discard sequence learning. To tackle these is…

Cited by 0SourceScholar
2018

TextSnake: A Flexible Representation for Detecting Text of Arbitrary Shapes

ECCV 2018poster

Driven by deep neural networks and large scale datasets, scene text detection methods have progressed substantially over the past years, continuously refreshing the performance records on various standard benchmarks. However, limited by the representations (axis-aligned rectangles, rotated rectangle…

Cited by 707SourcePDFScholar