← Search

Shai Mazor

10 accepted papers

2025

DocVLM: Make Your VLM an Efficient Reader

CVPR 2025poster

Vision-Language Models (VLMs) excel in diverse visual tasks but face challenges in document understanding, which requires fine-grained text processing. While typical visual tasks perform well with low-resolution inputs, reading-intensive applications demand high-resolution, resulting in significant…

Cited by 1SourcePDFScholar
2024

Question Aware Vision Transformer for Multimodal Reasoning

CVPR 2024highlight

Vision-Language (VL) models have gained significant research focus enabling remarkable advances in multimodal reasoning. These architectures typically comprise a vision encoder a Large Language Model (LLM) and a projection module that aligns visual features with the LLM's representation space. Despi…

Cited by 23SourcePDFScholar
2024

VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding

ECCV 2024poster

"In recent years, notable advancements have been made in the domain of visual document understanding, with the prevailing architecture comprising a cascade of vision and language models. The text component can either be extracted explicitly with the use of external OCR models in OCR-based approaches…

2023

CLIPTER: Looking at the Bigger Picture in Scene Text Recognition

ICCV 2023poster

Reading text in real-world scenarios often requires understanding the context surrounding it, especially when dealing with poor-quality text. However, current scene text recognizers are unaware of the bigger picture as they operate on cropped text images. In this study, we harness the representative…

Cited by 22PDFcodeScholar
2021

Sequence-to-Sequence Contrastive Learning for Text Recognition

CVPR 2021poster

We propose a framework for sequence-to-sequence contrastive learning (SeqCLR) of visual representations, which we apply to text recognition. To account for the sequence-to-sequence structure, each feature map is divided into different instances over which the contrastive loss is computed. This opera…

Cited by 162PDFcodeScholar
2020

Can You Read Me Now? Content Aware Rectification using Angle Supervision

ECCV 2020poster

The ubiquity of smartphone cameras has led to more and more documents being captured by cameras rather than scanned. Unlike flatbed scanners, photographed documents are often folded and crumpled, resulting in large local variance in text structure. The problem of document rectification is fundamenta…

Cited by 35SourcePDFScholar
2020

SCATTER: Selective Context Attentional Scene Text Recognizer

CVPR 2020poster

Scene Text Recognition (STR), the task of recognizing text against complex image backgrounds, is an active area of research. Current state-of-the-art (SOTA) methods still struggle to recognize text written in arbitrary shapes. In this paper, we introduce a novel architecture for STR, named Selective…

Cited by 194PDFcodeScholar
2020

ScrabbleGAN: Semi-Supervised Varying Length Handwritten Text Generation

CVPR 2020poster

Optical character recognition (OCR) systems performance have improved significantly in the deep learning era. This is especially true for handwritten text recognition (HTR), where each author has a unique style, unlike printed text, where the variation is smaller by design. That said, deep learning…

Cited by 179PDFScholar