← Search

Roy Ganz

9 accepted papers

2025

Adversaries With Incentives: A Strategic Alternative to Adversarial Robustness

ICLR 2025poster

Adversarial training aims to defend against *adversaries*: malicious opponents whose sole aim is to harm predictive performance in any way possible. This presents a rather harsh perspective, which we assert results in unnecessarily conservative training. As an alternative, we propose to model oppone…

2025

DocVLM: Make Your VLM an Efficient Reader

CVPR 2025poster

Vision-Language Models (VLMs) excel in diverse visual tasks but face challenges in document understanding, which requires fine-grained text processing. While typical visual tasks perform well with low-resolution inputs, reading-intensive applications demand high-resolution, resulting in significant…

Cited by 1SourcePDFScholar
2025

Paint by Inpaint: Learning to Add Image Objects by Removing Them First

CVPR 2025poster

Image editing has advanced significantly with the introduction of text-conditioned diffusion models. Despite this progress, seamlessly adding objects to images based on textual instructions without requiring user-provided input masks remains a challenge. We address this by leveraging the insight tha…

2024

Enhancing Consistency-Based Image Generation via Adversarialy-Trained Classification and Energy-Based Discrimination

NeurIPS 2024poster

The recently introduced Consistency models pose an efficient alternative to diffusion algorithms, enabling rapid and good quality image synthesis. These methods overcome the slowness of diffusion models by directly mapping noise to data, while maintaining a (relatively) simpler training. Consistency…

2024

GRAM: Global Reasoning for Multi-Page VQA

CVPR 2024poster

The increasing use of transformer-based large language models brings forward the challenge of processing long sequences. In document visual question answering (DocVQA) leading methods focus on the single-page setting while documents can span hundreds of pages. We present GRAM a method that seamlessl…

Cited by 12SourcePDFScholar
2024

Question Aware Vision Transformer for Multimodal Reasoning

CVPR 2024highlight

Vision-Language (VL) models have gained significant research focus enabling remarkable advances in multimodal reasoning. These architectures typically comprise a vision encoder a Large Language Model (LLM) and a projection module that aligns visual features with the LLM's representation space. Despi…

Cited by 23SourcePDFScholar
2023

CLIPTER: Looking at the Bigger Picture in Scene Text Recognition

ICCV 2023poster

Reading text in real-world scenarios often requires understanding the context surrounding it, especially when dealing with poor-quality text. However, current scene text recognizers are unaware of the bigger picture as they operate on cropped text images. In this study, we harness the representative…

Cited by 22PDFcodeScholar