← Search

Shai Bagon

13 accepted papers

2026

Hearing the Room Through the Shape of the Drum: Modal-Guided Sound Recovery from Multi-Point Surface Vibrations

CVPR 2026

Optical vibration sensing enables recovering the scene sound directly from the surface vibration of nearby objects, turning everyday objects into "visual microphones". However, most prior methods had focused on capturing the vibrations of specific objects with highly favorable vibration responses. T

Cited by 0SourcecodeScholar
2024

DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video

ECCV 2024poster

"We present – a new framework for long-term dense tracking in video. The pillar of our approach is combining test-time training on a single video, with the powerful localized semantic features learned by a pre-trained DINO-ViT model. Specifically, our framework simultaneously adopts DINO’s features…

Cited by 34SourcePDFScholar
2024

TokenFlow: Consistent Diffusion Features for Consistent Video Editing

ICLR 2024poster

The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual quality and user control over the generated content. In this work, we present a framework that harnesses the power of a text-to-i…

Cited by 244SourcePDFScholar
2023

Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation

CVPR 2023poster

Large-scale text-to-image generative models have been a revolutionary breakthrough in the evolution of generative AI, synthesizing diverse images with highly complex visual concepts. However, a pivotal challenge in leveraging such models for real-world content creation is providing users with contro…

2022

Combining Internal and External Constraints for Unrolling Shutter in Videos

ECCV 2022poster

"Videos obtained by rolling-shutter (RS) cameras result in spatially-distorted frames. These distortions become significant under fast camera/scene motions. Undoing effects of RS is sometimes addressed as a spatial problem, where objects need to be rectified/displaced in order to generate their corr…

Cited by 12SourcePDFScholar
2022

Diverse Generation from a Single Video Made Possible

ECCV 2022poster

"GANs are able to perform generation and manipulation tasks, trained on a single video. However, these single video GANs require unreasonable amount of time to train on a single video, rendering them almost impractical. In this paper we question the necessity of a GAN for generation from a single vi…

2022

Drop the GAN: In Defense of Patches Nearest Neighbors As Single Image Generative Models

CVPR 2022oral

Image manipulation dates back long before the deep learning era. The classical prevailing approaches were based on maximizing patch similarity between the input and generated output. Recently, single-image GANs were introduced as a superior and more sophisticated solution to image manipulation tasks…

Cited by 78PDFScholar
2021

Point of Care Image Analysis for COVID-19

ICASSP 2021accepted

Early detection of COVID-19 is key in containing the pandemic. Disease detection and evaluation based on imaging is fast and cheap and therefore plays an important role in COVID-19 handling. COVID-19 is easier to detect in chest CT, however, it is expensive, non-portable, and difficult to dis-infect…

Cited by 0SourceScholar
2020

Across Scales & Across Dimensions: Temporal Super-Resolution using Deep Internal Learning

ECCV 2020poster

When a very fast dynamic event is recorded with a low-framerate camera, the resulting video suffers from severe motion blur (due to exposure time) and motion aliasing (due to low sampling rate in time). True Temporal Super-Resolution (TSR) is more than just Temporal-Interpolation (increasing framera…