← Search

Zhineng Chen

29 accepted papers

2026

Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

AAAI 2026technical

Mathematical Expression Recognition (MER) has made significant progress in recognizing simple expressions, but the robust recognition of complex mathematical expressions with many tokens and multiple lines remains a formidable challenge. In this paper, we first introduce CMER-Bench, a carefully cons

Cited by 0SourcePDFScholar
2026

MDiff4STR: Mask Diffusion Model for Scene Text Recognition

AAAI 2026technical

Mask Diffusion Models (MDMs) have recently emerged as a promising alternative to auto-regressive models (ARMs) for vision-language tasks, owing to their flexible balance of efficiency and accuracy. In this paper, for the first time, we introduce MDMs into the Scene Text Recognition (STR) task. We sh

Cited by 0SourcePDFScholar
2026

SSR-SAM: Retrieval-Style Segment Anything Model for Semi-Supervised Ultra-High-Resolution Image Segmentation

AAAI 2026technical

Accurate segmentation of ultra-high-resolution (UHR) images, which often exceed tens of millions of pixels, is critically important in domains such as remote sensing and biomedical imaging. However, acquiring pixel-level annotations for such high-resolution images is prohibitively expensive and labo

Cited by 0SourcePDFScholar
2026

What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution

CVPR 2026

Large-scale and categorical-balanced text data is essential for training effective Scene Text Recognition (STR) models, which is hard to achieve when collecting real data. Synthetic data offers a cost-effective and perfectly labeled alternative. However, its performance often lags behind, revealing

Cited by 0SourcecodeScholar
2025

Confusion-Driven Self-Supervised Progressively Weighted Ensemble Learning for Non-Exemplar Class Incremental Learning

NeurIPS 2025poster

Non-exemplar class incremental learning (NECIL) aims to continuously assimilate new knowledge while retaining previously acquired knowledge in scenarios where prior examples are unavailable. A prevalent strategy within NECIL mitigates knowledge forgetting by freezing the feature extractor after trai…

Cited by 0SourceScholar
2025

Distilling Knowledge from Heterogeneous Architectures for Semantic Segmentation

AAAI 2025technical

Current knowledge distillation (KD) methods for semantic segmentation focus on guiding the student to imitate the teacher's knowledge within homogeneous architectures. However, these methods overlook the diverse knowledge contained in architectures with different inductive biases, which is crucial f…

Cited by 0SourcePDFScholar
2025

Explicit Relational Reasoning Network for Scene Text Detection

AAAI 2025technical

Connected component (CC) is a proper text shape representation that aligns with human reading intuition. However, CC-based text detection methods have recently faced a developmental bottleneck that their time-consuming post-processing is difficult to eliminate. To address this issue, we introduce an…

Cited by 0SourcePDFScholar
2025

IGD: Instructional Graphic Design with Multimodal Layer Generation

ICCV 2025poster

Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still labor-intensive. Existing diffusion-based methods generate non-editable g…

2025

Out of Length Text Recognition with Sub-String Matching

AAAI 2025technical

Scene Text Recognition (STR) methods have demonstrated robust performance in word-level text recognition. However, in real applications the text image is sometimes long due to detected with multiple horizontal words. It triggers the requirement to build long text recognition models from readily avai…

2025

SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition

ICCV 2025poster

Connectionist temporal classification (CTC)-based scene text recognition (STR) methods, e.g., SVTR, are widely employed in OCR applications, mainly due to their simple architecture, which only contains a visual model and a CTC-aligned linear classifier, and therefore fast inference. However, they ge…

2025

Stochasticity-aware No-Reference Point Cloud Quality Assessment

IJCAI 2025

The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) as a MOS regression problem and devise a deterministic mapping, ignoring the stochasticity in generating MOS from subject

Cited by 0SourcePDFScholar
2025

SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis

CVPR 2025poster

Due to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly.To address…

2025

TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition

ICCV 2025poster

Scene text recognition (STR) suffers from challenges of either less realistic synthetic training data or the difficulty of collecting sufficient high-quality real-world data, limiting the effectiveness of trained models. Meanwhile, despite producing holistically appealing text images, diffusion-base…

2024

DreamMesh: Jointly Manipulating and Texturing Triangle Meshes for Text-to-3D Generation

ECCV 2024poster

"Learning radiance fields (NeRF) with powerful 2D diffusion models has garnered popularity for text-to-3D generation. Nevertheless, the implicit 3D representations of NeRF lack explicit modeling of meshes and textures over surfaces, and such surface-undefined way may suffer from the issues, e.g., no…

2024

Dual Contrastive Learning Guided Pathological Image Re-Staining

ICASSP 2024accepted

Pathological virtual re-staining is a valuable research topic in AI-aided diagnosis, as it reduces the need for costly and time-consuming physical staining. However, existing methods still suffer from the insufficient ability to preserve tissue microstructure and cellular details, making the generat…

Cited by 0SourceScholar
2024

Improving Text-guided Object Inpainting with Semantic Pre-inpainting

ECCV 2024poster

"Recent years have witnessed the success of large text-to-image diffusion models and their remarkable potential to generate high-quality images. The further pursuit of enhancing the editability of images has sparked significant interest in the downstream task of inpainting a novel object described b…

2024

LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation Network

AAAI 2024technical

Recently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-…

2024

Learning to Rank Patches for Unbiased Image Redundancy Reduction

CVPR 2024poster

Images suffer from heavy spatial redundancy because pixels in neighboring regions are spatially correlated. Existing approaches strive to overcome this limitation by reducing less meaningful image regions. However current leading methods rely on supervisory signals. They may compel models to preserv…

2023

Bi-Directional Feature Fusion Generative Adversarial Network for Ultra-High Resolution Pathological Image Virtual Re-Staining

CVPR 2023poster

The cost of pathological examination makes virtual re-staining of pathological images meaningful. However, due to the ultra-high resolution of pathological images, traditional virtual re-staining methods have to divide a WSI image into patches for model training and inference. Such a limitation lead…

Cited by 11SourcePDFScholar
2023

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

ICCV 2023poster

Multilingual text recognition (MLTR) systems typically focus on a fixed set of languages, which makes it difficult to handle newly added languages or adapt to ever-changing data distribution. In this paper, we propose the Incremental MLTR (IMLTR) task in the context of incremental learning (IL), whe…

Cited by 17PDFcodeScholar
2023

Multi-Object Localization and Irrelevant-Semantic Separation for Nuclei Segmentation in Histopathology Images

ICASSP 2023accepted

Automated segmentation of nuclei in histopathology images is critical for cancer diagnosis and prognosis. Due to the high variability of nuclei morphology, numerous nuclei overlapping, and the wide existence of nuclei clusters, this task still remains challenging. In this paper, we propose an effect…

Cited by 0SourceScholar
2023

Prototypical Residual Networks for Anomaly Detection and Localization

CVPR 2023poster

Anomaly detection and localization are widely used in industrial manufacturing for its efficiency and effectiveness. Anomalies are rare and hard to collect and supervised models easily over-fit to these seen anomalies with a handful of abnormal samples, producing unsatisfactory performance. On the o…

Cited by 84SourcePDFScholar
2023

Resolving Task Confusion in Dynamic Expansion Architectures for Class Incremental Learning

AAAI 2023technical

The dynamic expansion architecture is becoming popular in class incremental learning, mainly due to its advantages in alleviating catastrophic forgetting. However, task confu- sion is not well assessed within this framework, e.g., the discrepancy between classes of different tasks is not well learne…

2023

TPS++: Attention-Enhanced Thin-Plate Spline for Scene Text Recognition

IJCAI 2023poster

Text irregularities pose significant challenges to scene text recognizers. Thin-Plate Spline (TPS)-based rectification is widely regarded as an effective means to deal with them. Currently, the calculation of TPS transformation parameters purely depends on the quality of regressed text borders. It i…

2022

Genre-Conditioned Long-Term 3D Dance Generation Driven by Music

ICASSP 2022accepted

Dancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre.…

Cited by 0SourceScholar
2022

SVTR: Scene Text Recognition with a Single Visual Model

IJCAI 2022poster

Dominant scene text recognition models commonly contain two building blocks, a visual model for feature extraction and a sequence model for text transcription. This hybrid architecture, although accurate, is complex and less efficient. In this study, we propose a Single Visual model for Scene Text r…

2021

A Hybrid Feature Enhancement Method for Gl And Segmentation In Histopathology Images

ICASSP 2021accepted

Accurate and automatic gland segmentation can help pathologists diagnose the malignancy of colorectal cancers. However, it remains a challenging task because of the large morphological differences between the glands and the presence of sticky glands. In this paper, a hybrid feature enhancement netwo…

Cited by 0SourceScholar
2017

Binarized Mode Seeking for Scalable Visual Pattern Discovery

CVPR 2017poster

This paper studies visual pattern discovery in large-scale image collections via binarized mode seeking, where images can only be represented as binary codes for efficient storage and computation. We address this problem from the perspective of binary space mode seeking. First, a binary mean shift (…

Cited by 10PDFScholar