← Search

Shan Guo

3 accepted papers

2025

InstructOCR: Instruction Boosting Scene Text Spotting

AAAI 2025technical

In the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene te…

2025

Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding

CVPR 2025poster

Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in docume…

2024

ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting

CVPR 2024poster

Abstract In recent years text-image joint pre-training techniques have shown promising results in various tasks. However in Optical Character Recognition (OCR) tasks aligning text instances with their corresponding text regions in images poses a challenge as it requires effective alignment between t…