← Search

Zhenrong Zhang

10 accepted papers

2026

THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning

ICLR 2026poster

Large Language Models (LLMs) have made remarkable progress in mathematical reasoning, but still continue to struggle with high-precision tasks like numerical computation and formal symbolic manipulation. Integrating external tools has emerged as a promising approach to bridge this gap. Despite recen…

Cited by 0SourcecodeScholar
2025

DocMamba: Efficient Document Pre-training with State Space Model

AAAI 2025technical

In recent years, visually-rich document understanding has attracted increasing attention. Transformer-based pre-trained models have become the mainstream approach, yielding significant performance gains in this field. However, the self-attention mechanism's quadratic computational complexity hinders…

2025

LP-Diff: Towards Improved Restoration of Real-World Degraded License Plate

CVPR 2025highlight

License plate (LP) recognition is crucial in intelligent traffic management systems. However, factors such as long distances and poor camera quality often lead to severe degradation of captured LP images, posing challenges to accurate recognition. The design of License Plate Image Restoration (LPIR)…

2025

RFL: Simplifying Chemical Structure Recognition with Ring-Free Language

AAAI 2025technical

The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional structures of molecules, particularly those with rings and multiple branches, present significant challenges for current…

2024

A Dataset and Model for Realistic License Plate Deblurring

IJCAI 2024poster

Vehicle license plate recognition is a crucial task in intelligent traffic management systems. However, the challenge of achieving accurate recognition persists due to motion blur from fast-moving vehicles. Despite the widespread use of image synthesis approaches in existing deblurring and recogniti…

2024

On the Federated Learning Framework for Cooperative Perception

RA-L 2024

Cooperative perception (CP) is essential to enhance the efficiency and safety of future transportation systems, requiring extensive data sharing among vehicles on the road, which raises significant privacy concerns. Federated learning offers a promising solution by enabling data privacy-preserving c

Cited by 10SourceScholar
2024

SEMv3: A Fast and Robust Approach to Table Separation Line Detection

IJCAI 2024poster

Table structure recognition (TSR) aims to parse the inherent structure of a table from its input image. The "split-and-merge" paradigm is a pivotal approach to parse table structure, where the table separation line detection is crucial. However, challenges such as wireless and deformed tables make i…

2024

UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition

EMNLP 2024finding

In the digital era, table structure recognition technology is a critical tool for processing and analyzing large volumes of tabular data. Previous methods primarily focus on visual aspects of table structure recovery but often fail to effectively comprehend the textual semantics within tables, parti…

2024

Viewing Writing as Video: Optical Flow based Multi-Modal Handwritten Mathematical Expression Recognition

ICASSP 2024accepted

Handwritten Mathematical Expression Recognition (HMER) forms a crucial task in the domain of document intelligence. It encompasses online and offline modalities, which utilize the trajectory sequence and static image as input, respectively. It is intuitive to utilize both online and offline modaliti…

Cited by 0SourceScholar
2023

HRDoc: Dataset and Baseline Method toward Hierarchical Reconstruction of Document Structures

AAAI 2023technical

The problem of document structure reconstruction refers to converting digital or scanned documents into corresponding semantic structures. Most existing works mainly focus on splitting the boundary of each element in a single document page, neglecting the reconstruction of semantic structure in mult…