← Search

Hadar Averbuch-Elor

27 accepted papers

2026

Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation

CVPR 2026

3D Gaussian Splatting has recently enabled fast and photorealistic reconstruction of static 3D scenes. However, dynamic editing of such scenes remains a significant challenge. We introduce a novel framework, Physics-Guided Score Distillation, to address a fundamental conflict: physics simulation pro

Cited by 0SourceScholar
2025

Blended Point Cloud Diffusion for Localized Text-guided Shape Editing

ICCV 2025poster

Natural language offers a highly intuitive interface for enabling localized fine-grained edits of 3D shapes. However, prior works face challenges in preserving global coherence while locally modifying the input 3D shape. In this work, we introduce an inpainting-based framework for editing shapes rep…

Cited by 0SourcePDFScholar
2025

ProtoSnap: Prototype Alignment For Cuneiform Signs

ICLR 2025poster

The cuneiform writing system served as the medium for transmitting knowledge in the ancient Near East for a period of over three thousand years. Cuneiform signs have a complex internal structure which is the subject of expert paleographic analysis, as variations in sign shapes bear witness to histor…

2025

WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild

NeurIPS 2025poster

Despite recent advances in sparse novel view synthesis (NVS) applied to object-centric scenes, scene-level NVS remains a challenge. A central issue is the lack of available clean multi-view training data, beyond manually curated datasets with limited diversity, camera variation, or licensing issues.…

Cited by 0SourceScholar
2024

ICC : Quantifying Image Caption Concreteness for Multimodal Dataset Curation

ACL 2024findings

Web-scale training on paired text-image data is becoming increasingly central to multimodal learning, but is challenged by the highly noisy nature of datasets in the wild. Standard data filtering approaches succeed in removing mismatched text-image pairs, but permit semantically related but highly a…

2024

Mitigating Open-Vocabulary Caption Hallucinations

EMNLP 2024main

While recent years have seen rapid progress in image-conditioned text generation, image captioning still suffers from the fundamental issue of hallucinations, namely, the generation of spurious details that cannot be inferred from the given image. Existing methods largely use closed-vocabulary objec…

2024

ReNoise: Real Image Inversion Through Iterative Noising

ECCV 2024poster

"Recent advancements in text-guided diffusion models have unlocked powerful image manipulation capabilities. However, applying these methods to real images necessitates the inversion of the images into the domain of the pretrained diffusion model. Achieving faithful inversion remains a challenge, pa…

Cited by 40SourcePDFScholar
2023

Doppelgangers: Learning to Disambiguate Images of Similar Structures

ICCV 2023oral

We consider the visual disambiguation task of determining whether a pair of visually similar images depict the same or distinct 3D surfaces (e.g., the same or opposite sides of a symmetric building). Illusory image matches, where two images observe distinct but visually similar 3D surfaces, can be c…

Cited by 37PDFcodeScholar
2023

Is BERT Blind? Exploring the Effect of Vision-and-Language Pretraining on Visual Language Understanding

CVPR 2023poster

Most humans use visual imagination to understand and reason about language, but models such as BERT reason about language using knowledge acquired during text-only pretraining. In this work, we investigate whether vision-and-language pretraining can improve performance on text-only tasks that involv…

2023

Localizing Object-Level Shape Variations with Text-to-Image Diffusion Models

ICCV 2023poster

Text-to-image models give rise to workflows which often begin with an exploration step, where users sift through a large collection of generated images. The global nature of the text-to-image generation process prevents users from narrowing their exploration to a particular object in the image. In t…

Cited by 123PDFScholar
2023

Neural Scene Chronology

CVPR 2023poster

In this work, we aim to reconstruct a time-varying 3D model, capable of rendering photo-realistic renderings with independent control of viewpoint, illumination, and time, from Internet photos of large-scale landmarks. The core challenges are twofold. First, different types of temporal changes, such…

2021

Extreme Rotation Estimation Using Dense Correlation Volumes

CVPR 2021poster

We present a technique for estimating the relative 3D rotation of an RGB image pair in an extreme setting, where the images have little or no overlap. We observe that, even when images do not overlap, there may be rich hidden cues as to their geometric relationship, such as light source directions,…

Cited by 48PDFcodeScholar
2021

Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision

ICCV 2021poster

The abundance and richness of Internet photos of landmarks and cities has led to significant progress in 3D vision over the past two decades, including automated 3D reconstructions of the world's landmarks from tourist photos. However, a major source of information available for these 3D-augmented c…

Cited by 20PDFcodeScholar
2021

Who's Waldo? Linking People Across Text and Images

ICCV 2021poster

We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which is predominantly object-based, our new task masks out the names of people in cap…

Cited by 22PDFcodeScholar
2020

DualSDF: Semantic Shape Manipulation Using a Two-Level Representation

CVPR 2020poster

We are seeing a Cambrian explosion of 3D shape representations for use in machine learning. Some representations seek high expressive power in capturing high-resolution detail. Other approaches seek to represent shapes as compositions of simple parts, which are intuitive for people to understand and…

Cited by 130PDFcodeScholar
2020

Hidden Footprints: Learning Contextual Walkability from 3D Human Trails

ECCV 2020poster

Predicting where people can walk in a scene is important for many tasks, including autonomous driving systems and human behavior analysis. Yet learning a computational model for this purpose is challenging due to semantic ambiguity and a lack of labeled data: current datasets only have labels on whe…

2020

Learning Gradient Fields for Shape Generation

ECCV 2020poster

In this work, we propose a novel technique to generate shapes from point cloud data. A point cloud can be viewed as samples from a distribution of 3D points whose density is concentrated near the surface of the shape. Point cloud generation thus amounts to moving randomly sampled points to high-dens…

2020

ScrabbleGAN: Semi-Supervised Varying Length Handwritten Text Generation

CVPR 2020poster

Optical character recognition (OCR) systems performance have improved significantly in the deep learning era. This is especially true for handwritten text recognition (HTR), where each author has a unique style, unlike printed text, where the variation is smaller by design. That said, deep learning…

Cited by 179PDFScholar