← Search

Zecheng Xie

6 accepted papers

2026

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

ICLR 2026poster

Recent advancements in multimodal slow-thinking systems have demonstrated remarkable performance across various visual reasoning tasks. However, their capabilities in text-rich image reasoning tasks remain understudied due to the absence of a dedicated and systematic benchmark. To address this gap,…

Cited by 0SourcecodeScholar
2025

DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

AAAI 2025technical

Current multimodal large language models (MLLMs) face significant challenges in visual document understanding (VDU) tasks due to the high resolution, dense text, and complex layouts typical of document images. These characteristics demand a high level of detail perception ability from MLLMs. While i…

Cited by 8SourcePDFScholar
2023

Improving Table Structure Recognition With Visual-Alignment Sequential Coordinate Modeling

CVPR 2023poster

Table structure recognition aims to extract the logical and physical structure of unstructured table images into a machine-readable format. The latest end-to-end image-to-text approaches simultaneously predict the two structures by two decoders, where the prediction of the physical structure (the bo…

Cited by 40SourcePDFScholar
2023

M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis

CVPR 2023poster

Document layout analysis is a crucial prerequisite for document understanding, including document retrieval and conversion. Most public datasets currently contain only PDF documents and lack realistic documents. Models trained on these datasets may not generalize well to real-world scenarios. Theref…

2019

Aggregation Cross-Entropy for Sequence Recognition

CVPR 2019oral

In this paper, we propose a novel method, aggregation cross-entropy (ACE), for sequence recognition from a brand new perspective. The ACE loss function exhibits competitive performance to CTC and the attention mechanism, with much quicker implementation (as it involves only four fundamental formulas…

Cited by 141PDFcodeScholar
2019

Tightness-Aware Evaluation Protocol for Scene Text Detection

CVPR 2019poster

Evaluation protocols play key role in the developmental progress of text detection methods. There are strict requirements to ensure that the evaluation methods are fair, objective and reasonable. However, existing metrics exhibit some obvious drawbacks: 1) They are not goal-oriented; 2) they cannot…

Cited by 42PDFcodeScholar