← Search

Yoad Tewel

7 accepted papers

2025

Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models

ICLR 2025poster

Adding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integrating the new object in a fitting location. Despite extensive efforts, existing models often struggle with this balance…

Cited by 5SourcePDFScholar
2025

Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models

ICLR 2025poster

Diffusion inversion is the problem of taking an image and a text prompt that describes it and finding a noise latent that would generate the exact same image. Most current deterministic inversion techniques operate by approximately solving an implicit equation and may converge slowly or yield poor…

2025

Make It Count: Text-to-Image Generation with an Accurate Number of Objects

CVPR 2025poster

Despite the unprecedented success of text-to-image diffusion models, controlling the number of depicted objects using text is surprisingly hard. This is important for various applications from technical documents, to children's books to illustrating cooking recipes. Generating object-correct counts…

Cited by 9SourcePDFScholar
2025

Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models

NAACL 2025long

Text-to-image (T2I) diffusion models rely on encoded prompts to guide the image generation process. Typically, these prompts are extended to a fixed length by appending padding tokens to the input. Despite being a default practice, the influence of padding tokens on the image generation process has…

Cited by 1SourcePDFScholar
2022

What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs

NeurIPS 2022accept

Given an input image, and nothing else, our method returns the bounding boxes of objects in the image and phrases that describe the objects. This is achieved within an open world paradigm, in which the objects in the input image may not have been encountered during the training of the localization m…

2022

ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

CVPR 2022poster

Recent text-to-image matching models apply contrastive learning to large corpora of uncurated pairs of images and sentences. While such models can provide a powerful score for matching and subsequent zero-shot tasks, they are not capable of generating caption given an image. In this work, we repurpo…

Cited by 185PDFcodeScholar