← Search

Shelly Sheynin

8 accepted papers

2025

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

CVPR 2025poster

We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into realistic video sequences based on a textual description. While recent advancements produce photorealistic outputs, they frequently struggle to create videos with accurate and consistent object mo…

2025

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

ICML 2025oral

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence.…

Cited by 8SourcePDFScholar
2024

Emu Edit: Precise Image Editing via Recognition and Generation Tasks

CVPR 2024highlight

Instruction-based image editing holds immense potential for a variety of applications as it enables users to perform any editing operation using a natural language instruction. However current models in this domain often struggle with accurately executing user instructions. We present Emu Edit a mul…

Cited by 124SourcePDFScholar
2024

Video Editing via Factorized Diffusion Distillation

ECCV 2024oral

"We introduce , a model that establishes a new state-of-the art in video editing without relying on any supervised video editing data. To develop we separately train an image editing adapter and a video generation adapter, and attach both to the same text-to-image model. Then, to align the adapters…

Cited by 12SourcePDFScholar
2023

Text-To-4D Dynamic Scene Generation

ICML 2023poster

We present MAV3D (Make-A-Video3D), a method for generating three-dimensional dynamic scenes from text descriptions. Our approach uses a 4D dynamic Neural Radiance Field (NeRF), which is optimized for scene appearance, density, and motion consistency by querying a Text-to-Video (T2V) diffusion-based…

2023

kNN-Diffusion: Image Generation via Large-Scale Retrieval

ICLR 2023poster

Recent text-to-image models have achieved impressive results. However, since they require large-scale datasets of text-image pairs, it is impractical to train them on new domains where data is scarce or not labeled. In this work, we propose using large-scale retrieval methods, in particular, efficie…

Cited by 139SourcePDFScholar
2022

Make-a-Scene: Scene-Based Text-to-Image Generation with Human Priors

ECCV 2022poster

"Recent text-to-image generation methods provide a simple yet exciting conversion capability between text and image domains. While these methods have incrementally improved the generated image fidelity and text relevancy, several pivotal gaps remain unanswered, limiting applicability and quality. We…

Cited by 575SourcePDFScholar
2021

A Hierarchical Transformation-Discriminating Generative Model for Few Shot Anomaly Detection

ICCV 2021poster

Anomaly detection, the task of identifying unusual samples in data, often relies on a large set of training samples. In this work, we consider the setting of few-shot anomaly detection in images, where only a few images are given at training. We devise a hierarchical generative model that captures t…

Cited by 109PDFScholar