← Search

Yingbin Zheng

6 accepted papers

2025

Achieving Ensemble-Like Performance in a Single Model: A Feature Diversification Framework for Image-Text Matching

AAAI 2025technical

Model ensembling is a widely used technique that enhances performance in image-text matching tasks by combining multiple models, each trained with different initializations. However, the inefficiencies associated with training several models and generating outputs from them constrain their practical…

Cited by 0SourcePDFScholar
2025

An Exemplar-based Framework for Chinese Text Recognition

AAAI 2025technical

This paper introduces a novel exemplar-based framework for reading Chinese texts in natural scene or document images. We present the Deep Exemplar-based Chinese Text Recognizer, which is structured to first identify candidate characters as exemplars from each text-line, and subsequently recognize th…

Cited by 0SourcePDFScholar
2025

Expanding the Scope of Negatives: Boosting Image-Text Matching with Negatives Distribution Guided Learning

AAAI 2025technical

Image-text matching is a crucial task that bridges visual and linguistic modalities. Recent research typically formulates it into the problem of maximizing the margin with the truly hardest negatives to enhance the learning efficiency and avoid the poor local optima. We argue that such formulation c…

Cited by 0SourcePDFScholar
2025

Unleashing the Semantic Adaptability of Controlled Diffusion Model for Image Colorization

IJCAI 2025

Recent data-driven image colorization methods have leveraged pre-trained Text-to-Image (T2I) diffusion models as generative prior, while still suffering from unsatisfactory and inaccurate semantic-level color control. To address these issues, we propose a Semantic Adaptation method (SeAda) that enha

2020

Scene Text Recognition with Temporal Convolutional Encoder

ICASSP 2020accepted

Texts from scene images typically consist of several characters and exhibit a characteristic sequence structure. Existing methods capture the structure with the sequence-to-sequence models by an encoder to have the visual representations and then a decoder to translate the features into the label se…

Cited by 0SourceScholar