← Search

Or Patashnik

21 accepted papers

2026

DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models

ICLR 2026poster

Text-to-Image (T2I) models have advanced rapidly, yet they remain vulnerable to semantic leakage, the unintended transfer of semantically related features between distinct entities. Existing mitigation strategies are often optimization-based or dependent on external inputs. We introduce **DeLeaker**…

Cited by 0SourceScholar
2026

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

CVPR 2026

Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language. Yet, relying solely on text instructions limits fine-grained control over the extent of edits. We introduce Kontinuous Kontext, an instruction-driven editing model that provides a new d

Cited by 0SourceScholar
2026

NoisePrints: Distortion-Free Watermarks for Authorship in Private Diffusion Models

ICLR 2026poster

With the rapid adoption of diffusion models for visual content generation, proving authorship and protecting copyright have become critical. This challenge is particularly important when model owners keep their models private and may be unwilling or unable to handle authorship issues, making third-p…

Cited by 0SourcecodeScholar
2026

Scaling Group Inference for Diverse and High-Quality Generation

ICLR 2026poster

Generative models typically sample outputs independently, and recent inference-time guidance and scaling algorithms focus on improving the quality of individual samples. However, in real-world applications, users are often presented with a set of multiple images (e.g., 4-8) for each prompt, where in…

Cited by 0SourcecodeScholar
2026

VLM-Guided Adaptive Negative Prompting for Creative Generation

ICLR 2026poster

Creative generation is the synthesis of new, surprising, and valuable samples that reflect user intent yet cannot be envisioned in advance. This task aims to extend human imagination, enabling the discovery of visual concepts that exist in the unexplored spaces between familiar domains. While text-t…

Cited by 0SourcecodeScholar
2026

Visual Diffusion Models are Geometric Solvers

CVPR 2026

In this paper we show that visual diffusion models can serve as effective geometric solvers: they can directly reason about geometric problems by working in pixel space. We first demonstrate this on the Inscribed Square Problem, a long-standing problem in geometry that asks whether every Jordan curv

Cited by 0SourceScholar
2025

Omni-ID: Holistic Identity Representation Designed for Generative Tasks

CVPR 2025poster

We introduce Omni-ID, a novel facial representation designed specifically for generative tasks. Omni-ID encodes holistic information about an individual's appearance across diverse expressions and poses within a fixed-size representation. It consolidates information from a varied number of unstructu…

Cited by 3SourcePDFScholar
2025

Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and Manipulation

CVPR 2025poster

Advancements in text-to-image diffusion models have led to significant progress in fast 3D content creation. One common approach is to generate a set of multi-view images of an object, and then reconstruct it into a 3D model. However, this approach bypasses the use of a native 3D representation of t…

Cited by 0SourcePDFScholar
2025

Stable Flow: Vital Layers for Training-Free Image Editing

CVPR 2025poster

Diffusion models have revolutionized the field of content synthesis and editing. Recent models have replaced the traditional UNet architecture with the Diffusion Transformer (DiT), and employed flow-matching for improved training and sampling. However, they exhibit limited generation diversity. In t…

2024

Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation

ECCV 2024poster

"Text-to-image diffusion models have an unprecedented ability to generate diverse and high-quality images. However, they often struggle to faithfully capture the intended semantics of complex input prompts that include multiple subjects. Recently, numerous layout-to-image extensions have been introd…

Cited by 26SourcePDFScholar
2024

CLiC: Concept Learning in Context

CVPR 2024highlight

This paper addresses the challenge of learning a local visual pattern of an object from one image and generating images depicting objects with that pattern. Learning a localized concept and placing it on an object in a target image is a nontrivial task as the objects may have different orientations…

Cited by 17SourcePDFScholar
2024

LCM-Lookahead for Encoder-based Text-to-Image Personalization

ECCV 2024poster

"Recent advancements in diffusion models have introduced fast sampling methods that can effectively produce high-quality images in just one or a few denoising steps. Interestingly, when these are distilled from existing diffusion models, they often maintain alignment with the original model, retaini…

2024

ReNoise: Real Image Inversion Through Iterative Noising

ECCV 2024poster

"Recent advancements in text-guided diffusion models have unlocked powerful image manipulation capabilities. However, applying these methods to real images necessitates the inversion of the images into the domain of the pretrained diffusion model. Achieving faithful inversion remains a challenge, pa…

Cited by 40SourcePDFScholar
2023

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

ICLR 2023top-25%

Text-to-image models offer unprecedented freedom to guide creation through natural language. Yet, it is unclear how such freedom can be exercised to generate images of specific unique concepts, modify their appearance, or compose them in new roles and novel scenes. In other words, we ask: how can we…

2023

Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures

CVPR 2023poster

Text-guided image generation has progressed rapidly in recent years, inspiring major breakthroughs in text-guided shape generation. Recently, it has been shown that using score distillation, one can successfully text-guide a NeRF model to generate a 3D object. We adapt the score distillation to the…

2023

Localizing Object-Level Shape Variations with Text-to-Image Diffusion Models

ICCV 2023poster

Text-to-image models give rise to workflows which often begin with an exploration step, where users sift through a large collection of generated images. The global nature of the text-to-image generation process prevents users from narrowing their exploration to a particular object in the image. In t…

Cited by 123PDFScholar
2021

Encoding in Style: A StyleGAN Encoder for Image-to-Image Translation

CVPR 2021poster

We present a generic image-to-image translation framework, pixel2style2pixel (pSp). Our pSp framework is based on a novel encoder network that directly generates a series of style vectors which are fed into a pretrained StyleGAN generator, forming the extended W+ latent space. We first show that our…

Cited by 1395PDFcodeScholar
2021

StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery

ICCV 2021poster

Inspired by the ability of StyleGAN to generate highly re-alistic images in a variety of domains, much recent work hasfocused on understanding how to use the latent spaces ofStyleGAN to manipulate generated and real images. How-ever, discovering semantically meaningful latent manipula-tions typicall…

Cited by 1379PDFcodeScholar