← Search

Sagie Benaim

25 accepted papers

2026

Designing a Conditional Prior Distribution for Flow-Based Generative Models

ICML 2026poster

Flow-based generative models have recently shown impressive performance for conditional generation tasks, such as text-to-image generation. However, current methods transform a general unimodal noise distribution to a specific mode of the target data distribution. As such, every point in the initial…

Cited by 9SourcecodeScholar
2026

DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion

ICML 2026poster

Diffusion Transformer models can generate images with remarkable fidelity and detail, yet training them at ultra-high resolutions remains extremely costly due to the self-attention mechanism's quadratic scaling with the number of image tokens. In this paper, we introduce Dynamic Position Extrapolati…

Cited by 0SourceScholar
2026

Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation

CVPR 2026

3D Gaussian Splatting has recently enabled fast and photorealistic reconstruction of static 3D scenes. However, dynamic editing of such scenes remains a significant challenge. We introduce a novel framework, Physics-Guided Score Distillation, to address a fundamental conflict: physics simulation pro

Cited by 0SourceScholar
2026

Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware Distillation

ICLR 2026poster

Vision Foundation Models (VFMs) have achieved remarkable success when applied to various downstream 2D tasks. Despite their effectiveness, they often exhibit a critical lack of 3D awareness. To this end, we introduce Splat and Distill, a framework that instills robust 3D awareness into 2D VFMs by au…

Cited by 0SourcecodeScholar
2025

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

CVPR 2025poster

We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into realistic video sequences based on a textual description. While recent advancements produce photorealistic outputs, they frequently struggle to create videos with accurate and consistent object mo…

2024

Coarse-To-Fine Tensor Trains for Compact Visual Representations

ICML 2024poster

The ability to learn compact, high-quality, and easy-to-optimize representations for visual data is paramount to many applications such as novel view synthesis and 3D reconstruction. Recent work has shown substantial success in using tensor networks to design such compact and high-quality representa…

2024

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

AAAI 2024technical

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio: globally, the input audio is semantically associated with t…

2023

Discriminative Class Tokens for Text-to-Image Diffusion Models

ICCV 2023poster

Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in the input text. One way of alleviating these issues is to trai…

Cited by 10PDFcodeScholar
2023

Reward Finetuning for Faster and More Accurate Unsupervised Object Discovery

NeurIPS 2023poster

Recent advances in machine learning have shown that Reinforcement Learning from Human Feedback (RLHF) can improve machine learning models and align them with human preferences. Although very successful for Large Language Models (LLMs), these advancements have not had a comparable impact in research…

2022

Polynomial Neural Fields for Subband Decomposition and Manipulation

NeurIPS 2022accept

Neural fields have emerged as a new paradigm for representing signals, thanks to their ability to do it compactly while being easy to optimize. In most applications, however, neural fields are treated like a black box, which precludes many signal manipulation tasks. In this paper, we propose a new c…

2022

Text2Mesh: Text-Driven Neural Stylization for Meshes

CVPR 2022oral

In this work, we develop intuitive controls for editing the style of 3D objects. Our framework, Text2Mesh, stylizes a 3D mesh by predicting color and local geometric details which conform to a target text prompt. We consider a disentangled representation of a 3D object using a fixed mesh input (cont…

Cited by 386PDFcodeScholar
2021

A Hierarchical Transformation-Discriminating Generative Model for Few Shot Anomaly Detection

ICCV 2021poster

Anomaly detection, the task of identifying unusual samples in data, often relies on a large set of training samples. In this work, we consider the setting of few-shot anomaly detection in images, where only a few images are given at training. We devise a hierarchical generative model that captures t…

Cited by 109PDFScholar
2021

Permuted AdaIN: Reducing the Bias Towards Global Statistics in Image Classification

CVPR 2021poster

Recent work has shown that convolutional neural network classifiers overly rely on texture at the expense of shape cues. We make a similar but different distinction between shape and local image cues, on the one hand, and global image statistics, on the other. Our method, called Permuted Adaptive In…

Cited by 115PDFcodeScholar
2020

Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single Sample

NeurIPS 2020poster

We consider the task of generating diverse and novel videos from a single video sample. Recently, new hierarchical patch-GAN based approaches were proposed for generating diverse images, given only a single sample at training time. Moving to videos, these approaches fail to generate diverse samples,…

2020

SpeedNet: Learning the Speediness in Videos

CVPR 2020oral

We wish to automatically predict the "speediness" of moving objects in videos - whether they move faster, at, or slower than their "natural" speed. The core component in our approach is SpeedNet--a novel deep network trained to detect if a video is playing at normal rate, or if it is sped up. SpeedN…

Cited by 322PDFScholar
2019

Emerging Disentanglement in Auto-Encoder Based Unsupervised Image Content Transfer

ICLR 2019poster

We study the problem of learning to map, in an unsupervised way, between domains $A$ and $B$, such that the samples $\vb \in B$ contain all the information that exists in samples $\va\in A$ and some additional information. For example, ignoring occlusions, $B$ can be people with glasses, $A$ people…

Cited by 46SourcePDFScholar
2019

Semi-supervised Monaural Singing Voice Separation with a Masking Network Trained on Synthetic Mixtures

ICASSP 2019accepted

We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers th…

Cited by 0SourceScholar
2018

Estimating the Success of Unsupervised Image to Image Translation

ECCV 2018poster

While in supervised learning, the validation error is an unbiased estimator of the generalization (test) error and complexity-based generalization bounds are abundant, no such bounds exist for learning a mapping in an unsupervised way. As a result, when training GANs and specifically when using GANs…

2018

The Role of Minimal Complexity Functions in Unsupervised Learning of Semantic Mappings

ICLR 2018poster

We discuss the feasibility of the following learning problem: given unmatched samples from two domains and nothing else, learn a mapping between the two, which preserves semantics. Due to the lack of paired samples and without any definition of the semantic information, the problem might seem ill-po…

Cited by 25SourcePDFScholar