Vision Mamba-Based Approach for Incomplete Boundary Document Image Rectification
Weihao Zhang, Xin Xia, Maopeng Li, Yunbo Zhao
Abstract
Capturing document images using handheld mobile devices often results in geometric deformations, which adversely affect the accuracy of Optical Character Recognition (OCR) and document understanding. However, existing transformer-based methods face significant computational costs when processing document images on resource-constrained devices. This study proposes an enhanced Vision Mamba architecture to learn the structural information of document images, thereby rectifying deformed images while reducing computational resource consumption. Additionally, owing to the relative positioning between the document and the imaging device, captured images may exhibit incomplete boundaries. Conventional learning-based methods are primarily designed for images with complete boundaries, which can diminish correction effectiveness. To address this issue, mask consistency loss and preprocessing techniques are introduced to improve the rectification of document images with incomplete boundaries. Experimental results demonstrate the effectiveness and superiority of this method, highlighting its significant value for intelligent document processing.
BibTeX
@inproceedings{icassp2025_visionmambabased,
title = {Vision Mamba-Based Approach for Incomplete Boundary Document Image Rectification},
author = {Weihao Zhang and Xin Xia and Maopeng Li and Yunbo Zhao},
booktitle = {ICASSP 2025},
year = {2025}
}