← Search

Alona Golts

3 accepted papers

2025

DocVLM: Make Your VLM an Efficient Reader

CVPR 2025poster

Vision-Language Models (VLMs) excel in diverse visual tasks but face challenges in document understanding, which requires fine-grained text processing. While typical visual tasks perform well with low-resolution inputs, reading-intensive applications demand high-resolution, resulting in significant…

Cited by 1SourcePDFScholar
2024

GRAM: Global Reasoning for Multi-Page VQA

CVPR 2024poster

The increasing use of transformer-based large language models brings forward the challenge of processing long sequences. In document visual question answering (DocVQA) leading methods focus on the single-page setting while documents can span hundreds of pages. We present GRAM a method that seamlessl…

Cited by 12SourcePDFScholar
2023

CLIPTER: Looking at the Bigger Picture in Scene Text Recognition

ICCV 2023poster

Reading text in real-world scenarios often requires understanding the context surrounding it, especially when dealing with poor-quality text. However, current scene text recognizers are unaware of the bigger picture as they operate on cropped text images. In this study, we harness the representative…

Cited by 22PDFcodeScholar