← Search

Rinon Gal

9 accepted papers

2026

ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation

ICLR 2026poster

While recent generative models synthesize high-quality visual content, they still struggle with generating rare or fine-grained concepts. To address this challenge, we explore the usage of Retrieval-Augmented Generation (RAG) for image generation, and introduce ImageRAG, a training-free method for r…

Cited by 0SourcecodeScholar
2025

Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models

ICLR 2025poster

Adding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integrating the new object in a fitting location. Despite extensive efforts, existing models often struggle with this balance…

Cited by 5SourcePDFScholar
2025

Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models

NAACL 2025long

Text-to-image (T2I) diffusion models rely on encoded prompts to guide the image generation process. Typically, these prompts are extended to a fixed length by appending padding tokens to the input. Despite being a default practice, the influence of padding tokens on the image generation process has…

Cited by 1SourcePDFScholar
2024

Breathing Life Into Sketches Using Text-to-Video Priors

CVPR 2024highlight

A sketch is one of the most intuitive and versatile tools humans use to convey their ideas visually. An animated sketch opens another dimension to the expression of ideas and is widely used by designers for a variety of purposes. Animating sketches is a laborious process requiring extensive experien…

Cited by 29SourcePDFScholar
2023

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

ICLR 2023top-25%

Text-to-image models offer unprecedented freedom to guide creation through natural language. Yet, it is unclear how such freedom can be exercised to generate images of specific unique concepts, modify their appearance, or compose them in new roles and novel scenes. In other words, we ask: how can we…

2022

"“This Is My Unicorn, Fluffy”: Personalizing Frozen Vision-Language Representations"

ECCV 2022poster

"Large Vision & Language models pretrained on web-scale data provide representations that are invaluable for numerous V&L problems. However, it is unclear how they can be extended to reason about user-specific visual concepts in unstructured language. This problem arises in multiple domains, from pe…

Cited by 94SourcePDFScholar
2022

HyperStyle: StyleGAN Inversion With HyperNetworks for Real Image Editing

CVPR 2022poster

The inversion of real images into StyleGAN's latent space is a well-studied problem. Nevertheless, applying existing approaches to real-world scenarios remains an open challenge, due to an inherent trade-off between reconstruction and editability: latent space regions which can accurately represent…

Cited by 330PDFcodeScholar