← Search

Tali Dekel

27 accepted papers

2025

Generative Omnimatte: Learning to Decompose Video into Layers

CVPR 2025highlight

Given a video and a set of input object masks, an omnimatte method aims to decompose the video into semantically meaningful layers containing individual objects along with their associated effects, such as shadows and reflections.Existing omnimatte methods assume a static background or accurate pose…

Cited by 3SourcePDFScholar
2024

DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video

ECCV 2024poster

"We present – a new framework for long-term dense tracking in video. The pillar of our approach is combining test-time training on a single video, with the powerful localized semantic features learned by a pre-trained DINO-ViT model. Specifically, our framework simultaneously adopts DINO’s features…

Cited by 34SourcePDFScholar
2024

Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer

CVPR 2024poster

We present a new method for text-driven motion transfer - synthesizing a video that complies with an input text prompt describing the target objects and scene while maintaining an input video's motion and scene layout. Prior methods are confined to transferring motion across two subjects within the…

Cited by 42SourcePDFScholar
2024

TokenFlow: Consistent Diffusion Features for Consistent Video Editing

ICLR 2024poster

The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual quality and user control over the generated content. In this work, we present a framework that harnesses the power of a text-to-i…

Cited by 244SourcePDFScholar
2023

DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

NeurIPS 2023spotlight

Current perceptual similarity metrics operate at the level of pixels and patches. These metrics compare images in terms of their low-level colors and textures, but fail to capture mid-level similarities and differences in image layout, object pose, and semantic content. In this paper, we develop a p…

2023

Imagic: Text-Based Real Image Editing With Diffusion Models

CVPR 2023poster

Text-conditioned image editing has recently attracted considerable interest. However, most methods are currently limited to one of the following: specific editing types (e.g., object overlay, style transfer), synthetically generated images, or requiring multiple input images of a common object. In t…

Cited by 1151SourcePDFScholar
2023

MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

ICML 2023poster

Recent advances in text-to-image generation with diffusion models present transformative capabilities in image quality. However, user controllability of the generated image, and fast adaptation to new tasks still remains an open challenge, currently mostly addressed by costly and long re-training an…

2023

Neural Congealing: Aligning Images to a Joint Semantic Atlas

CVPR 2023poster

We present Neural Congealing -- a zero-shot self-supervised framework for detecting and jointly aligning semantically-common content across a given set of images. Our approach harnesses the power of pre-trained DINO-ViT features to learn: (i) a joint semantic atlas -- a 2D grid that captures the mod…

Cited by 16SourcePDFScholar
2023

Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation

CVPR 2023poster

Large-scale text-to-image generative models have been a revolutionary breakthrough in the evolution of generative AI, synthesizing diverse images with highly complex visual concepts. However, a pivotal challenge in leveraging such models for real-world content creation is providing users with contro…

2022

Associating Objects and Their Effects in Video through Coordination Games

NeurIPS 2022accept

We explore a feed-forward approach for decomposing a video into layers, where each layer contains an object of interest along with its associated shadows, reflections, and other visual effects. This problem is challenging since associated effects vary widely with the 3D geometry and lighting conditi…

Cited by 5SourcePDFScholar
2022

Diverse Generation from a Single Video Made Possible

ECCV 2022poster

"GANs are able to perform generation and manipulation tasks, trained on a single video. However, these single video GANs require unreasonable amount of time to train on a single video, rendering them almost impractical. In this paper we question the necessity of a GAN for generation from a single vi…

2022

Text2LIVE: Text-Driven Layered Image and Video Editing

ECCV 2022poster

"We present a method for zero-shot, text-driven editing of natural images and videos. Given an image or a video and a text prompt, our goal is to edit the appearance of existing objects (e.g., texture) or augment the scene with visual effects (e.g., smoke, fire) in a semantic manner. We train a gene…

2021

Omnimatte: Associating Objects and Their Effects in Video

CVPR 2021poster

Computer vision has become increasingly better at segmenting objects in images and videos; however, scene effects related to the objects -- shadows, reflections, generated smoke, etc. -- are typically overlooked. Identifying such scene effects and associating them with the objects producing them is…

Cited by 55PDFScholar
2020

Semantic Pyramid for Image Generation

CVPR 2020oral

We present a novel GAN-based model that utilizes the space of deep features learned by a pre-trained classification model. Inspired by classical image pyramid representations, we construct our model as a Semantic Generation Pyramid -- a hierarchical framework which leverages the continuum of semanti…

Cited by 66PDFScholar
2020

SpeedNet: Learning the Speediness in Videos

CVPR 2020oral

We wish to automatically predict the "speediness" of moving objects in videos - whether they move faster, at, or slower than their "natural" speed. The core component in our approach is SpeedNet--a novel deep network trained to detect if a video is playing at normal rate, or if it is sped up. SpeedN…

Cited by 322PDFScholar
2019

Learning the Depths of Moving People by Watching Frozen People

CVPR 2019oral

We present a method for predicting dense depth in scenarios where both a monocular camera and people in the scene are freely moving. Existing methods for recovering depth for dynamic, non-rigid objects from monocular video impose strong assumptions on the objects' motion and may only recover sparse…

Cited by 276PDFScholar
2019

Speech2Face: Learning the Face Behind a Voice

CVPR 2019poster

How much can we infer about a person's looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking. We design and train a deep neural network to perform this task using millions of natural Internet/…

Cited by 222PDFcodeScholar
2018

Sparse, Smart Contours to Represent and Edit Images

CVPR 2018poster

We study the problem of reconstructing an image from information stored at contour locations. We show that high-quality reconstructions with high fidelity to the source image can be obtained from sparse input, e.g., comprising less than 6% of image pixels. This is a significant improvement over exis…

Cited by 96SourcePDFScholar
2015

Best-Buddies Similarity for Robust Template Matching

CVPR 2015poster

We propose a novel method for template matching in unconstrained environments. Its essence is the Best Buddies Similarity (BBS), a useful, robust, and parameter-free similarity measure between two sets of points. BBS is based on a count of Best Buddies Pairs (BBPs)--pairs of points in which each one…

Cited by 205SourcePDFScholar