← Search

Lior Wolf

131 accepted papers

2025

Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds

CVPR 2025poster

In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their training data. This limitation significantly impedes their c…

2025

Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models

ICLR 2025poster

Adding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integrating the new object in a fitting location. Despite extensive efforts, existing models often struggle with this balance…

Cited by 5SourcePDFScholar
2025

DeciMamba: Exploring the Length Extrapolation Potential of Mamba

ICLR 2025poster

Long-range sequence processing poses a significant challenge for Transformers due to their quadratic complexity in input length. A promising alternative is Mamba, which demonstrates high performance and achieves Transformer-level capabilities while requiring substantially fewer computational resourc…

2025

Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation

ICLR 2025poster

Recent advances in efficient sequence modeling have led to attention-free layers, such as Mamba, RWKV, and various gated RNNs, all featuring sub-quadratic complexity in sequence length and excellent scaling properties, enabling the construction of a new type of foundation models. In this paper, we p…

2025

FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

NeurIPS 2025poster

Text-to-video diffusion models are notoriously limited in their ability to model temporal aspects such as motion, physics, and dynamic interactions. Existing approaches address this limitation by retraining the model or introducing external conditioning signals to enforce temporal consistency. In th…

Cited by 0SourceScholar
2025

Predicting local fMRI activations from EEG: a Feasibility Study Using Both Classical and Modern Machine Learning Pipelines

ICASSP 2025accepted

fMRI’s clinical use is limited by cost, while EEG is more accessible but lacks spatial detail and deep brain coverage. Research aims to predict deep brain activations from combined fMRI and EEG data. We compare classical machine learning and a CNN-transformer pipeline for this mapping across multipl…

Cited by 0SourceScholar
2025

Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability

NeurIPS 2025poster

The development of effective explainability tools for Transformers is a crucial pursuit in deep learning research. One of the most promising approaches in this domain is Layer-wise Relevance Propagation (LRP), which propagates relevance scores backward through the network to the input space by redis…

Cited by 0SourceScholar
2025

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

ICML 2025oral

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence.…

Cited by 8SourcePDFScholar
2024

Anomaly Detection with Variance Stabilized Density Estimation

UAI 2024poster

We propose a modified density estimation problem that is highly effective for detecting anomalies in tabular data. Our approach assumes that the density function is relatively stable (with lower variance) around normal samples. We have verified this hypothesis empirically using a wide range of real-…

2024

Backward Lens: Projecting Language Model Gradients into the Vocabulary Space

EMNLP 2024main

Understanding how Transformer-based Language Models (LMs) learn and recall information is a key goal of the deep learning community. Recent interpretability methods project weights and hidden states obtained from the forward pass to the models’ vocabularies, helping to uncover how information flows…

2024

Converting Transformers to Polynomial Form for Secure Inference Over Homomorphic Encryption

ICML 2024poster

Designing privacy-preserving DL solutions is a major challenge within the AI community. Homomorphic Encryption (HE) has emerged as one of the most promising approaches in this realm, enabling the decoupling of knowledge between a model owner and a data owner. Despite extensive research and applicati…

Cited by 24SourcePDFScholar
2024

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

AAAI 2024technical

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio: globally, the input audio is semantically associated with t…

2024

Knowledge Editing in Language Models via Adapted Direct Preference Optimization

EMNLP 2024finding

Large Language Models (LLMs) can become outdated over time as they may lack updated world knowledge, leading to factual knowledge errors and gaps. Knowledge Editing (KE) aims to overcome this challenge using weight updates that do not require expensive retraining. We propose treating KE as an LLM al…

Cited by 9SourcePDFScholar
2024

Separate and Diffuse: Using a Pretrained Diffusion Model for Better Source Separation

ICLR 2024poster

The problem of speech separation, also known as the cocktail party problem, refers to the task of isolating a single speech signal from a mixture of speech signals. Previous work on source separation derived an upper bound for the source separation task in the domain of human speech. This bound is d…

Cited by 7SourcePDFScholar
2024

The Hidden Language of Diffusion Models

ICLR 2024poster

Text-to-image diffusion models have demonstrated an unparalleled ability to generate high-quality, diverse images from a textual prompt. However, the internal representations learned by these models remain an enigma. In this work, we present Conceptor, a novel method to interpret the internal repres…

2023

Box-based Refinement for Weakly Supervised and Unsupervised Localization Tasks

ICCV 2023poster

It has been established that training a box-based detector network can enhance the localization performance of weakly supervised and unsupervised methods. Moreover, we extend this understanding by demonstrating that these detectors can be utilized to improve the original network, paving the way for…

Cited by 8PDFcodeScholar
2023

Decision S4: Efficient Sequence-Based RL via State Spaces Layers

ICLR 2023poster

Recently, sequence learning methods have been applied to the problem of off-policy Reinforcement Learning, including the seminal work on Decision Transformers, which employs transformers for this task. Since transformers are parameter-heavy, cannot benefit from history longer than a fixed window siz…

Cited by 30SourcePDFScholar
2023

Discriminative Class Tokens for Text-to-Image Diffusion Models

ICCV 2023poster

Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in the input text. One way of alleviating these issues is to trai…

Cited by 10PDFcodeScholar
2023

Semi-supervised learning of partial differential operators and dynamical flows

UAI 2023poster

The evolution of many dynamical systems is generically governed by nonlinear partial differential equations (PDEs), whose solution, in a simulation framework, requires vast amounts of computational resources. In this work, we present a novel method that combines a hyper-network solver with a Fourier…

2022

Dynamic Dual-Output Diffusion Models

CVPR 2022poster

Iterative denoising-based generation, also known as denoising diffusion models, has recently been shown to be comparable in quality to other classes of generative models, and even surpass them. Including, in particular, Generative Adversarial Networks, which are currently the state of the art in man…

Cited by 34PDFScholar
2022

No Token Left Behind: Explainability-Aided Image Classification and Generation

ECCV 2022poster

"The application of zero-shot learning in computer vision has been revolutionized by the use of image-text matching models. The most notable example, CLIP, has been widely used for both zero-shot classification and guiding generative models with a text prompt. However, the zero-shot use of CLIP is u…

2022

Optimizing Relevance Maps of Vision Transformers Improves Robustness

NeurIPS 2022accept

It has been observed that visual classification models often rely mostly on spurious cues such as the image background, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and direct the model to base its predictio…

2022

Unsupervised Disentanglement with Tensor Product Representations on the Torus

ICLR 2022poster

The current methods for learning representations with auto-encoders almost exclusively employ vectors as the latent representations. In this work, we propose to employ a tensor product structure for this purpose. This way, the obtained representations are naturally disentangled. In contrast to the…

2022

What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs

NeurIPS 2022accept

Given an input image, and nothing else, our method returns the bounding boxes of objects in the image and phrases that describe the objects. This is achieved within an open world paradigm, in which the objects in the input image may not have been encountered during the training of the localization m…

2022

XAI for Transformers: Better Explanations through Conservative Propagation

ICML 2022spotlight

Transformers have become an important workhorse of machine learning, with numerous applications. This necessitates the development of reliable methods for increasing their transparency. Multiple interpretability methods, often based on gradient information, have been proposed. We show that the gradi…

2022

ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

CVPR 2022poster

Recent text-to-image matching models apply contrastive learning to large corpora of uncurated pairs of images and sentences. While such models can provide a powerful score for matching and subsequent zero-shot tasks, they are not capable of generating caption given an image. In this work, we repurpo…

Cited by 185PDFcodeScholar
2021

A Hierarchical Transformation-Discriminating Generative Model for Few Shot Anomaly Detection

ICCV 2021poster

Anomaly detection, the task of identifying unusual samples in data, often relies on a large set of training samples. In this work, we consider the setting of few-shot anomaly detection in images, where only a few images are given at training. We devise a hierarchical generative model that captures t…

Cited by 109PDFScholar
2021

Generic Attention-Model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers

ICCV 2021poster

Transformers are increasingly dominating multi-modal reasoning tasks, such as visual question answering, achieving state-of-the-art results thanks to their ability to contextualize information using the self-attention and co-attention mechanisms. These attention modules also play a role in other com…

Cited by 392PDFcodeScholar
2021

High Fidelity Speech Regeneration with Application to Speech Enhancement

ICASSP 2021accepted

Speech enhancement has seen great improvement in recent years mainly through contributions in denoising, speaker separation, and dereverberation methods that mostly deal with environmental effects on vocal audio. To enhance speech beyond the limitations of the original signal, we take a regeneration…

Cited by 0SourceScholar
2021

Permuted AdaIN: Reducing the Bias Towards Global Statistics in Image Classification

CVPR 2021poster

Recent work has shown that convolutional neural network classifiers overly rely on texture at the expense of shape cues. We make a similar but different distinction between shape and local image cues, on the one hand, and global image statistics, on the other. Our method, called Permuted Adaptive In…

Cited by 115PDFcodeScholar
2021

Single Channel Voice Separation for Unknown Number of Speakers Under Reverberant and Noisy Settings

ICASSP 2021accepted

We present a unified network for voice separation of an unknown number of speakers. The proposed approach is composed of several separation heads optimized together with a speaker classification branch. The separation is carried out in the time domain, together with parameter sharing between all sep…

Cited by 0SourceScholar
2021

Visualization of Supervised and Self-Supervised Neural Networks via Attribution Guided Factorization

AAAI 2021technical

Neural network visualization techniques mark image locations by their relevancy to the network's classification. Existing methods are effective in highlighting the regions that affect the resulting classification the most. However, as we show, these methods are limited in their ability to identify t…

2020

Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single Sample

NeurIPS 2020poster

We consider the task of generating diverse and novel videos from a single video sample. Recently, new hierarchical patch-GAN based approaches were proposed for generating diverse images, given only a single sample at training time. Moving to videos, these approaches fail to generate diverse samples,…

2020

OneGAN: Simultaneous Unsupervised Learning of Conditional Image Generation, Foreground Segmentation, and Fine-Grained Clustering

ECCV 2020poster

Foreground Segmentation, and Fine-Grained Clustering","We present a method for simultaneously learning, in an unsupervised manner, (i) a conditional image generator, (ii) foreground extraction and segmentation, (iii) clustering into a two-level class hierarchy, and (iv) object removal and background…

Cited by 48SourcePDFScholar
2019

Emerging Disentanglement in Auto-Encoder Based Unsupervised Image Content Transfer

ICLR 2019poster

We study the problem of learning to map, in an unsupervised way, between domains $A$ and $B$, such that the samples $\vb \in B$ contain all the information that exists in samples $\va\in A$ and some additional information. For example, ignoring occlusions, $B$ can be people with glasses, $A$ people…

Cited by 46SourcePDFScholar
2019

Semi-supervised Monaural Singing Voice Separation with a Masking Network Trained on Synthetic Mixtures

ICASSP 2019accepted

We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers th…

Cited by 0SourceScholar
2019

Unsupervised Microvascular Image Segmentation Using an Active Contours Mimicking Neural Network

ICCV 2019accepted

The task of blood vessel segmentation in microscopy images is crucial for many diagnostic and research applications. However, vessels can look vastly different, depending on the transient imaging conditions, and collecting data for supervised training is laborious. We present a novel deep learning m…

2018

Estimating the Success of Unsupervised Image to Image Translation

ECCV 2018poster

While in supervised learning, the validation error is an unbiased estimator of the generalization (test) error and complexity-based generalization bounds are abundant, no such bounds exist for learning a mapping in an unsupervised way. As a result, when training GANs and specifically when using GANs…

2018

The Role of Minimal Complexity Functions in Unsupervised Learning of Semantic Mappings

ICLR 2018poster

We discuss the feasibility of the following learning problem: given unmatched samples from two domains and nothing else, learn a mapping between the two, which preserves semantics. Due to the lack of paired samples and without any definition of the semantic information, the problem might seem ill-po…

Cited by 25SourcePDFScholar
2018

VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop

ICLR 2018poster

We present a new neural text to speech (TTS) method that is able to transform text to speech in voices that are sampled in the wild. Unlike other systems, our solution is able to deal with unconstrained voice samples and without requiring aligned phonemes or linguistic features. The network architec…

2015

Associating Neural Word Embeddings With Deep Image Representations Using Fisher Vectors

CVPR 2015poster

In recent years, the problem of associating a sentence with an image has gained a lot of attention. This work continues to push the envelope and makes further progress in the performance of image annotation and image search by a sentence tasks. In this work, we are using the Fisher Vector as a sente…

Cited by 354SourcePDFScholar
2015

Live Repetition Counting

ICCV 2015poster

The task of counting the number of repetitions of approximately the same action in an input video sequence is addressed. The proposed method runs online and not on the complete pre-captured video. It analyzes sequentially blocks of 20 non-consecutive frames. The cycle length within each block is eva…

Cited by 111PDFScholar