← Search

Kurban Ubul

7 accepted papers

2025

ASANet: Scene Text Recognition With Alternate Self-Attention

ICASSP 2025accepted

Text recognition in complex scenes is a challenging task in Computer Vision. In this paper, we propose an innovative framework for scene text recognition, ASANet, which features an Alternate Attention Enhancement Encoder and a Masked Dual-modal Decoder. The encoder incorporates a 12-layer Alternatin…

Cited by 0SourceScholar
2025

GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts

ICCV 2025poster

Low-light enhancement has wide applications in autonomous driving, 3D reconstruction, remote sensing, surveillance, and so on, which can significantly improve information utilization. However, most existing methods lack generalization and are limited to specific tasks such as image recovery. To addr…

2025

KCGAFormer: When Large-Kernel ConvFormer Meets KAN in Semantic Segmentation

ICASSP 2025accepted

Vision Transformer, with the distinctive architecture and self-attention mechanisms, had profoundly influenced the field of computer vision, establishing Transformer-based models as benchmarks for semantic segmentation. In this study, we propose a pioneering hybrid model that fuses Kolmogorov-Arnold…

Cited by 0SourceScholar
2025

Long-tailed Oracle Character Recognition Based on Convolutional Neural Networks and Vision Transformers

ICASSP 2025accepted

Oracle bone inscriptions, which are among the oldest known hieroglyphics in China, encompass rich historical and cultural information. However, the automatic recognition of oracle characters faces substantial challenges due to issues with data quality and long-tail distribution. This study introduce…

Cited by 0SourceScholar
2024

MMHSV: A Multimodal Handwritten Signature Verification Fusing Dynamic and Static Feature

ICASSP 2024accepted

In recent years, significant progress has been made in the field of handwritten signature verification through methods based on deep learning. However, due to the high intra-class variability and high inter-class similarity of signature samples, achieving high accuracy and security in handwritten si…

Cited by 0SourceScholar
2024

The Collaboration of 3D Convolutions and CRO-TSM in Lipreading

ICASSP 2024accepted

Lip reading refers to the recognition of speech solely based on the subtle movements of the lips without audio information. Extracting temporal information in lip reading has always been a challenge in this field. In this work, we propose an effective method for extracting temporal information. Spec…

Cited by 0SourceScholar
2021

How to Use Time Information Effectively? Combining with Time Shift Module for Lipreading

ICASSP 2021accepted

Lipreading refers to recognizing the speaker's speech content through the image sequence of lip movement without the speech signal. Currently, most models use a spatiotemporal (3D) convolutional layer combined with 2D CNN to extract spatial and temporal features from image sequences. However, compar…

Cited by 0SourceScholar