← Search

Idan Schwartz

13 accepted papers

2024

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

AAAI 2024technical

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio: globally, the input audio is semantically associated with t…

2023

Discriminative Class Tokens for Text-to-Image Diffusion Models

ICCV 2023poster

Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in the input text. One way of alleviating these issues is to trai…

Cited by 10PDFcodeScholar
2022

Optimizing Relevance Maps of Vision Transformers Improves Robustness

NeurIPS 2022accept

It has been observed that visual classification models often rely mostly on spurious cues such as the image background, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and direct the model to base its predictio…

2022

ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

CVPR 2022poster

Recent text-to-image matching models apply contrastive learning to large corpora of uncurated pairs of images and sentences. While such models can provide a powerful score for matching and subsequent zero-shot tasks, they are not capable of generating caption given an image. In this work, we repurpo…

Cited by 185PDFcodeScholar
2021

Perceptual Score: What Data Modalities Does Your Model Perceive?

NeurIPS 2021poster

Machine learning advances in the last decade have relied significantly on large-scale datasets that continue to grow in size. Increasingly, those datasets also contain different data modalities. However, large multi-modal datasets are hard to annotate, and annotations may contain biases that we are…

2020

Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies

NeurIPS 2020poster

Many recent datasets contain a variety of different data modalities, for instance, image, question, and answer data in visual question answering (VQA). When training deep net classifiers on those multi-modal datasets, the modalities get exploited at different scales, i.e., some modalities can more e…

2017

High-Order Attention Models for Visual Question Answering

NeurIPS 2017poster

The quest for algorithms that enable cognitive abilities is an important part of machine learning. A common trait in many recently investigated cognitive-like tasks is that they take into account different data modalities, such as visual and textual input. In this paper we propose a novel and gene…