← Search

Amir Hertz

9 accepted papers

2024

Curved Diffusion: A Generative Model With Optical Geometry Control

ECCV 2024poster

"State-of-the-art diffusion models can generate highly realistic images based on various conditioning like text, segmentation, and depth. However, an essential aspect often overlooked is the specific camera geometry used during image capture. The influence of different optical systems on the final s…

Cited by 1SourcePDFScholar
2024

Style Aligned Image Generation via Shared Attention

CVPR 2024poster

Large-scale Text-to-Image (T2I) models have rapidly gained prominence across creative fields generating visually compelling outputs from textual prompts. However controlling these models to ensure consistent style remains challenging with existing methods necessitating fine-tuning and manual interve…

2023

Delta Denoising Score

ICCV 2023poster

This paper introduces Delta Denoising Score (DDS), a novel diffusion-based scoring technique that optimizes a parametric model for the task of image editing. Unlike the existing Score Distillation Sampling (SDS), which queries the generative model with a single image-text pair, DDS utilizes an addit…

Cited by 119PDFScholar
2023

NULL-Text Inversion for Editing Real Images Using Guided Diffusion Models

CVPR 2023poster

Recent large-scale text-guided diffusion models provide powerful image generation capabilities. Currently, a massive effort is given to enable the modification of these images using text only as means to offer intuitive and versatile editing tools. To edit a real image using these state-of-the-art t…

2023

Prompt-to-Prompt Image Editing with Cross-Attention Control

ICLR 2023top-25%

Recent large-scale text-driven synthesis diffusion models have attracted much attention thanks to their remarkable capabilities of generating highly diverse images that follow given text prompts. Therefore, it is only natural to build upon these synthesis models to provide text-driven image editing…

2022

MotionCLIP: Exposing Human Motion Generation to CLIP Space

ECCV 2022poster

"We introduce MotionCLIP, a 3D human motion auto-encoder featuring a latent embedding that is disentangled, well behaved, and supports highly semantic textual descriptions. MotionCLIP gains its unique power by aligning its latent space with that of the Contrastive Language-Image Pre-training (CLIP)…

2021

SAPE: Spatially-Adaptive Progressive Encoding for Neural Optimization

NeurIPS 2021poster

Multilayer-perceptrons (MLP) are known to struggle learning functions of high-frequencies, and in particular, instances of wide frequency bands. We present a progressive mapping scheme for input signals of MLP networks, enabling them to better fit a wide range of frequencies without sacrificing tra…

Cited by 70SourcePDFScholar