← Search

Daniel Cohen-or

59 accepted papers

2026

Cycle-Consistent Tuning for Layered Image Decomposition

CVPR 2026

Disentangling visual layers in real-world images is a persistent challenge in vision and graphics, as such layers often involve non-linear and globally coupled interactions, including shading, reflection, and perspective distortion. In this work, we present an in-context image decomposition framewor

Cited by 0SourcecodeScholar
2026

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

CVPR 2026

Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language. Yet, relying solely on text instructions limits fine-grained control over the extent of edits. We introduce Kontinuous Kontext, an instruction-driven editing model that provides a new d

Cited by 0SourceScholar
2026

Visual Diffusion Models are Geometric Solvers

CVPR 2026

In this paper we show that visual diffusion models can serve as effective geometric solvers: they can directly reason about geometric problems by working in pixel space. We first demonstrate this on the Inscribed Square Problem, a long-standing problem in geometry that asks whether every Jordan curv

Cited by 0SourceScholar
2025

Dynamic Typography: Bringing Text to Life via Video Diffusion Prior

ICCV 2025poster

Text animation serves as an expressive medium, transforming static communication into dynamic experiences by infusing words with motion to evoke emotions, emphasize meanings, and construct compelling narratives. Crafting animations that are semantically aware poses significant challenges, demanding…

Cited by 0SourcePDFScholar
2025

EmoEdit: Evoking Emotions through Image Manipulation

CVPR 2025poster

Affective Image Manipulation (AIM) seeks to modify user-provided images to evoke specific emotions. This task is inherently complex due to its twofold objective: evoking the intended emotion while preserving image composition. Existing AIM methods primarily adjust color and style, often failing to e…

2025

Layered Image Vectorization via Semantic Simplification

CVPR 2025poster

This work presents a progressive image vectorization technique that reconstructs the raster image as layer-wise vectors from semantic-aligned macro structures to finer details. Our approach introduces a new image simplification method leveraging the feature-average effect in the Score Distillation S…

Cited by 6SourcePDFScholar
2025

NeuralSVG: An Implicit Representation for Text-to-Vector Generation

ICCV 2025poster

Vector graphics are essential in design, providing artists with a versatile medium for creating resolution-independent and highly editable visual content. Recent advancements in vision-language and diffusion models have fueled interest in text-to-vector graphics generation. However, existing approac…

2025

Omni-ID: Holistic Identity Representation Designed for Generative Tasks

CVPR 2025poster

We introduce Omni-ID, a novel facial representation designed specifically for generative tasks. Omni-ID encodes holistic information about an individual's appearance across diverse expressions and poses within a fixed-size representation. It consolidates information from a varied number of unstructu…

Cited by 3SourcePDFScholar
2025

Stable Flow: Vital Layers for Training-Free Image Editing

CVPR 2025poster

Diffusion models have revolutionized the field of content synthesis and editing. Recent models have replaced the traditional UNet architecture with the Diffusion Transformer (DiT), and employed flow-matching for improved training and sampling. However, they exhibit limited generation diversity. In t…

2025

TriTex: Learning Texture from a Single Mesh via Triplane Semantic Features

CVPR 2025poster

As 3D content creation continues to grow, transferring semantic textures between 3D meshes remains a significant challenge in computer graphics. While recent methods leverage text-to-image diffusion models for texturing, they often struggle to preserve the appearance of the source texture during tex…

Cited by 0SourcePDFScholar
2024

Breathing Life Into Sketches Using Text-to-Video Priors

CVPR 2024highlight

A sketch is one of the most intuitive and versatile tools humans use to convey their ideas visually. An animated sketch opens another dimension to the expression of ideas and is widely used by designers for a variety of purposes. Animating sketches is a laborious process requiring extensive experien…

Cited by 29SourcePDFScholar
2024

CLiC: Concept Learning in Context

CVPR 2024highlight

This paper addresses the challenge of learning a local visual pattern of an object from one image and generating images depicting objects with that pattern. Learning a localized concept and placing it on an object in a target image is a nontrivial task as the objects may have different orientations…

Cited by 17SourcePDFScholar
2024

Curved Diffusion: A Generative Model With Optical Geometry Control

ECCV 2024poster

"State-of-the-art diffusion models can generate highly realistic images based on various conditioning like text, segmentation, and depth. However, an essential aspect often overlooked is the specific camera geometry used during image capture. The influence of different optical systems on the final s…

Cited by 1SourcePDFScholar
2024

Generating Non-Stationary Textures using Self-Rectification

CVPR 2024poster

This paper addresses the challenge of example-based non-stationary texture synthesis. We introduce a novel two-step approach wherein users first modify a reference texture using standard image editing tools yielding an initial rough target for the synthesis. Subsequently our proposed method termed "…

2024

MAS: Multi-view Ancestral Sampling for 3D Motion Generation Using 2D Diffusion

CVPR 2024poster

We introduce Multi-view Ancestral Sampling (MAS) a method for 3D motion generation using 2D diffusion models that were trained on motions obtained from in-the-wild videos. As such MAS opens opportunities to exciting and diverse fields of motion previously under-explored as 3D data is scarce and hard…

2024

Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment

ECCV 2024poster

"While existing image-text alignment models reach high quality binary assessments, they fall short of pinpointing the exact source of misalignment. In this paper, we present a method to provide detailed textual and visual explanation of detected misalignments between text-image pairs. We leverage la…

2024

SAI3D: Segment Any Instance in 3D Scenes

CVPR 2024poster

Advancements in 3D instance segmentation have traditionally been tethered to the availability of annotated datasets limiting their application to a narrow spectrum of object categories. Recent efforts have sought to harness vision-language models like CLIP for open-set semantic reasoning yet these m…

2024

Single Motion Diffusion

ICLR 2024spotlight

Synthesizing realistic animations of humans, animals, and even imaginary creatures, has long been a goal for artists and computer graphics professionals. Compared to the imaging domain, which is rich with large available datasets, the number of data instances for the motion domain is limited, partic…

2024

Style Aligned Image Generation via Shared Attention

CVPR 2024poster

Large-scale Text-to-Image (T2I) models have rapidly gained prominence across creative fields generating visually compelling outputs from textual prompts. However controlling these models to ensure consistent style remains challenging with existing methods necessitating fine-tuning and manual interve…

2023

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

ICLR 2023top-25%

Text-to-image models offer unprecedented freedom to guide creation through natural language. Yet, it is unclear how such freedom can be exercised to generate images of specific unique concepts, modify their appearance, or compose them in new roles and novel scenes. In other words, we ask: how can we…

2023

CLIPascene: Scene Sketching with Different Types and Levels of Abstraction

ICCV 2023oral

In this paper, we present a method for converting a given scene image into a sketch using different types and multiple levels of abstraction. We distinguish between two types of abstraction. The first considers the fidelity of the sketch, varying its representation from a more precise portrayal of…

Cited by 101PDFScholar
2023

Delta Denoising Score

ICCV 2023poster

This paper introduces Delta Denoising Score (DDS), a novel diffusion-based scoring technique that optimizes a parametric model for the task of image editing. Unlike the existing Score Distillation Sampling (SDS), which queries the generative model with a single image-text pair, DDS utilizes an addit…

Cited by 119PDFScholar
2023

Domain Expansion of Image Generators

CVPR 2023poster

Can one inject new concepts into an already trained generative model, while respecting its existing structure and knowledge? We propose a new task -- domain expansion -- to address this. Given a pretrained generator and novel (but related) domains, we expand the generator to jointly model all domain…

Cited by 17SourcePDFScholar
2023

Human Motion Diffusion Model

ICLR 2023top-25%

Natural and expressive human motion generation is the holy grail of computer animation. It is a challenging task, due to the diversity of possible motion, human perceptual sensitivity to it, and the difficulty of accurately describing it. Therefore, current generative solutions are either low-qualit…

2023

Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures

CVPR 2023poster

Text-guided image generation has progressed rapidly in recent years, inspiring major breakthroughs in text-guided shape generation. Recently, it has been shown that using score distillation, one can successfully text-guide a NeRF model to generate a 3D object. We adapt the score distillation to the…

2023

Localizing Object-Level Shape Variations with Text-to-Image Diffusion Models

ICCV 2023poster

Text-to-image models give rise to workflows which often begin with an exploration step, where users sift through a large collection of generated images. The global nature of the text-to-image generation process prevents users from narrowing their exploration to a particular object in the image. In t…

Cited by 123PDFScholar
2023

MoDi: Unconditional Motion Synthesis From Diverse Data

CVPR 2023poster

The emergence of neural networks has revolutionized the field of motion synthesis. Yet, learning to unconditionally synthesize motions from a given distribution remains challenging, especially when the motions are highly diverse. In this work, we present MoDi -- a generative model trained in an unsu…

2023

NULL-Text Inversion for Editing Real Images Using Guided Diffusion Models

CVPR 2023poster

Recent large-scale text-guided diffusion models provide powerful image generation capabilities. Currently, a massive effort is given to enable the modification of these images using text only as means to offer intuitive and versatile editing tools. To edit a real image using these state-of-the-art t…

2023

Prompt-to-Prompt Image Editing with Cross-Attention Control

ICLR 2023top-25%

Recent large-scale text-driven synthesis diffusion models have attracted much attention thanks to their remarkable capabilities of generating highly diverse images that follow given text prompts. Therefore, it is only natural to build upon these synthesis models to provide text-driven image editing…

2023

SKED: Sketch-guided Text-based 3D Editing

ICCV 2023poster

Text-to-image diffusion models are gradually introduced into computer graphics, recently enabling the development of Text-to-3D pipelines in an open domain. However, for interactive editing purposes, local manipulations of content through a simplistic textual interface can be arduous. Incorporating…

Cited by 73PDFcodeScholar
2022

FLEX: Extrinsic Parameters-Free Multi-View 3D Human Motion Reconstruction

ECCV 2022poster

"The increasing availability of video recordings made by multiple cameras has offered new means for mitigating occlusion and depth ambiguities in pose and motion reconstruction methods. Yet, multi-view algorithms strongly depend on camera parameters, particularly on relative transformations between…

2022

MotionCLIP: Exposing Human Motion Generation to CLIP Space

ECCV 2022poster

"We introduce MotionCLIP, a 3D human motion auto-encoder featuring a latent embedding that is disentangled, well behaved, and supports highly semantic textual descriptions. MotionCLIP gains its unique power by aligning its latent space with that of the Contrastive Language-Image Pre-training (CLIP)…

2022

ShapeFormer: Transformer-Based Shape Completion via Sparse Representation

CVPR 2022poster

We present ShapeFormer, a transformer-based network that produces a distribution of object completions, conditioned on incomplete, and possibly noisy, point clouds. The resultant distribution can then be sampled to generate likely completions, each of which exhibits plausible shape details, while be…

Cited by 160PDFScholar
2021

Encoding in Style: A StyleGAN Encoder for Image-to-Image Translation

CVPR 2021poster

We present a generic image-to-image translation framework, pixel2style2pixel (pSp). Our pSp framework is based on a novel encoder network that directly generates a series of style vectors which are fed into a pretrained StyleGAN generator, forming the extended W+ latent space. We first show that our…

Cited by 1395PDFcodeScholar
2021

SAPE: Spatially-Adaptive Progressive Encoding for Neural Optimization

NeurIPS 2021poster

Multilayer-perceptrons (MLP) are known to struggle learning functions of high-frequencies, and in particular, instances of wide frequency bands. We present a progressive mapping scheme for input signals of MLP networks, enabling them to better fit a wide range of frequencies without sacrificing tra…

Cited by 70SourcePDFScholar
2021

ShapeConv: Shape-Aware Convolutional Layer for Indoor RGB-D Semantic Segmentation

ICCV 2021poster

RGB-D semantic segmentation has attracted increasing attention over the past few years. Existing methods mostly employ homogeneous convolution operators to consume the RGB and depth features, ignoring their intrinsic differences. In fact, the RGB values capture the photometric appearance properties…

Cited by 199PDFcodeScholar
2021

StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery

ICCV 2021poster

Inspired by the ability of StyleGAN to generate highly re-alistic images in a variety of domains, much recent work hasfocused on understanding how to use the latent spaces ofStyleGAN to manipulate generated and real images. How-ever, discovering semantically meaningful latent manipula-tions typicall…

Cited by 1379PDFcodeScholar
2020

GANHopper: Multi-Hop GAN for Unsupervised Image-to-Image Translation

ECCV 2020poster

We introduce GANHopper, an unsupervised image-to-image translation network that transforms images gradually between two domains, through multiple hops. Instead of executing translation directly, we steer the translation by requiring the network to produce in-between images that resemble weighted hyb…

Cited by 31SourcePDFScholar
2020

Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation

CVPR 2020poster

Many applications, such as autonomous driving, heavily rely on multi-modal data where spatial alignment between the modalities is required. Most multi-modal registration methods struggle computing the spatial correspondence between the images using prevalent cross-modality similarity measures. In th…

Cited by 180PDFScholar
2019

CompoNet: Learning to Generate the Unseen by Part Synthesis and Composition

ICCV 2019poster

Data-driven generative modeling has made remarkable progress by leveraging the power of deep neural networks. A reoccurring challenge is how to enable a model to generate a rich variety of samples from the entire target distribution, rather than only from a distribution confined to the training data…

Cited by 64PDFcodeScholar
2019

Patch-Based Progressive 3D Point Set Upsampling

CVPR 2019poster

We present a detail-driven deep neural network for point set upsampling. A high-resolution point set is essential for point-based rendering and surface reconstruction. Inspired by the recent success of neural image super-resolution techniques, we progressively train a cascade of patch-based upsampli…

Cited by 353PDFcodeScholar
2019

ZigZagNet: Fusing Top-Down and Bottom-Up Context for Object Segmentation

CVPR 2019poster

Multi-scale context information has proven to be essential for object segmentation tasks. Recent works construct the multi-scale context by aggregating convolutional feature maps extracted by different levels of a deep neural network. This is typically done by propagating and fusing features in a on…

Cited by 86PDFcodeScholar
2018

EC-Net: an Edge-aware Point set Consolidation Network

ECCV 2018poster

Point clouds obtained from 3D scans are typically sparse, irregular, and noisy, and required to be consolidated. In this paper, we present the first deep learning based {em edge-aware} technique to facilitate the consolidation of point clouds. We design our network to process points grouped in local…

Cited by 340SourcePDFScholar
2018

Multi-Scale Context Intertwining for Semantic Segmentation

ECCV 2018poster

Accurate semantic image segmentation requires the joint consideration of local appearance, semantic information, and global scene context. In today’s age of pre-trained deep networks and their powerful convolutional features, state-of-the-art semantic segmentation approaches differ mostly in how the…

Cited by 211SourcePDFScholar
2018

PU-Net: Point Cloud Upsampling Network

CVPR 2018poster

Learning and analyzing 3D point clouds with deep networks is challenging due to the sparseness and irregularity of the data. In this paper, we present a data-driven point cloud upsampling technique. The key idea is to learn multi-level features per point and expand the point set via a multi-branch c…

2018

Specular-to-Diffuse Translation for Multi-View Reconstruction

ECCV 2018poster

Most multi-view 3D reconstruction algorithms, especially when shape-from-shading cues are used, assume that object appearance is predominantly diffuse. To alleviate this restriction, we introduce S2Dnet, a generative adversarial network for transferring multiple views of objects with specular reflec…

Cited by 28SourcePDFScholar
2017

Cascaded Feature Network for Semantic Segmentation of RGB-D Images

ICCV 2017poster

Fully convolutional network (FCN) has been successfully applied in semantic segmentation of scenes represented with RGB images. Images augmented with depth channel provide more understanding of the geometric information of the scene in the image. The question is how to best exploit this additional i…

Cited by 175PDFScholar