← Search

Chengquan Zhang

14 accepted papers

2026

MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation

CVPR 2026

End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding. Despite advances in vision-language large models (VLLMs), robustness across diverse visual scenes and low-resource langu

Cited by 0SourceScholar
2026

Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity

ICLR 2026poster

Vision-language models (VLMs) face significant computational inefficiencies caused by excessive generation of visual tokens. While prior work shows that a large fraction of visual tokens are redundant, existing compression methods struggle to balance \textit{importance preservation} and \textit{info…

Cited by 0SourcecodeScholar
2026

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

CVPR 2026

Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Traditional cascaded pipelines depend on precise layout analysis and often fail under casually captured or non-standard conditions. Although end-to-end approa

Cited by 0SourceScholar
2026

Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs

ICLR 2026poster

Direct Preference Optimization (DPO) has emerged as a cornerstone of reinforcement learning from human feedback (RLHF) due to its simplicity and efficiency. However, existing DPO-based methods typically treat all preference pairs equally, overlooking substantial variations in data quality and learni…

Cited by 0SourceScholar
2024

Textual Grounding for Open-vocabulary Visual Information Extraction in Layout-diversified Documents

ECCV 2024poster

"Current methodologies have achieved notable success in the closed-set visual information extraction (VIE) task, while the exploration into open-vocabulary settings is comparatively underdeveloped, which is practical for individual users in terms of inferring information across documents of diverse…

Cited by 1SourcePDFScholar
2024

Towards Unified Multi-granularity Text Detection with Interactive Attention

ICML 2024spotlight

Existing OCR engines or document image analysis systems typically rely on training separate models for text detection in varying scenarios and granularities, leading to significant computational complexity and resource demands. In this paper, we introduce "Detect Any Text" (DAT), an advanced paradig…

Cited by 1SourcePDFScholar
2024

WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-only Supervised Text Spotting

ECCV 2024poster

"Transcription-only Supervised Text Spotting aims to learn text spotters relying only on transcriptions but no text boundaries for supervision, thus eliminating expensive boundary annotation. The crux of this task lies in locating each transcription in scene text images without location annotations.…

2023

Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding

IJCAI 2023poster

Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are challenging to be directly adapted to model document. They are unabl…

Cited by 8SourcePDFScholar
2023

StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

ICLR 2023poster

In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-training tasks: masked image modeling and masked language modeling, based on text region-level image masking. The proposed…

2021

PGNet: Real-time Arbitrarily-Shaped Text Spotting with Point Gathering Network

AAAI 2021technical

The reading of arbitrarily-shaped text has received increasing research attention. However, existing text spotters are mostly built on two-stage frameworks or character-based methods, which suffer from either Non-Maximum Suppression (NMS), Region-of-Interest (RoI) operations, or character-level anno…

2020

Towards Accurate Scene Text Recognition With Semantic Reasoning Networks

CVPR 2020poster

Scene text image contains two levels of contents: visual texture and semantic information. Although the previous scene text recognition methods have made great progress over the past few years, the research on mining semantic information to assist text recognition attracts less attention, only RNN-l…

Cited by 424PDFScholar
2019

Look More Than Once: An Accurate Detector for Text of Arbitrary Shapes

CVPR 2019poster

Previous scene text detection methods have progressed substantially over the past years. However, limited by the receptive field of CNNs and the simple representations like rectangle bounding box or quadrangle adopted to describe text, previous methods may fall short when dealing with more challengi…

Cited by 328PDFScholar
2017

WordSup: Exploiting Word Annotations for Character Based Text Detection

ICCV 2017poster

Imagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese, Japanese, mathematical expression and etc. It is natural and conven…

Cited by 250PDFScholar
2016

Multi-Oriented Text Detection With Fully Convolutional Networks

CVPR 2016poster

In this paper, we propose an unconventional approach for text detection in natural images. Both global and local cues are taken into account for localizing text lines in a coarse-to-fine procedure. First, a Fully Convolutional Network (FCN) model is trained for predicting a salient map of text regio…

Cited by 740PDFScholar