← Search

Umapada Pal

6 accepted papers

2024

Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes

ICRA 2024poster

When used in a real-world noisy environment, the capacity to generalize to multiple domains is essential for any autonomous scene text spotting system. However, existing state-of-the-art methods employ pretraining and fine-tuning strategies on natural scene datasets, which do not exploit the feature…

Cited by 9SourceScholar
2024

Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network

AAAI 2024technical

Scene text recognition is inherently a vision-language task. However, previous works have predominantly focused either on extracting more robust visual features or designing better language modeling. How to effectively and jointly model vision and language to mitigate heavy reliance on a single moda…

2022

SGBANet: Semantic GAN and Balanced Attention Network for Arbitrarily Oriented Scene Text Recognition

ECCV 2022poster

"Scene text recognition is a challenging task due to the complex backgrounds and diverse variations of text instances. In this paper, we propose a novel Semantic GAN and Balanced Attention Network (SGBANet) to recognize the texts in scene images. The proposed method first generates the simple semant…

Cited by 32SourcePDFScholar
2022

TIPS: Text-Induced Pose Synthesis

ECCV 2022poster

"In computer vision, human pose synthesis and transfer deal with probabilistic image generation of a person in a previously unseen pose from an already available observation of that person. Though researchers have recently proposed several methods to achieve this task, most of these techniques deriv…

2021

LoOp: Looking for Optimal Hard Negative Embeddings for Deep Metric Learning

ICCV 2021poster

Deep metric learning has been effectively used to learn distance metrics for different visual tasks like image retrieval, clustering, etc. In order to aid the training process, existing methods either use a hard mining strategy to extract the most informative samples or seek to generate hard synthet…

Cited by 23PDFcodeScholar
2020

STEFANN: Scene Text Editor Using Font Adaptive Neural Network

CVPR 2020poster

Textual information in a captured scene plays an important role in scene interpretation and decision making. Though there exist methods that can successfully detect and interpret complex text regions present in a scene, to the best of our knowledge, there is no significant prior work that aims to mo…

Cited by 91PDFcodeScholar