← Search

Hila Chefer

11 accepted papers

2026

Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

ICML 2026poster

Strong semantic representations improve the convergence and generation quality of diffusion and flow models. Existing approaches largely rely on external models, which require separate training, operate on misaligned objectives, and exhibit unexpected scaling behavior. We argue that this dependence …

Cited by 0SourceScholar
2025

FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

NeurIPS 2025poster

Text-to-video diffusion models are notoriously limited in their ability to model temporal aspects such as motion, physics, and dynamic interactions. Existing approaches address this limitation by retraining the model or introducing external conditioning signals to enforce temporal consistency. In th…

Cited by 0SourceScholar
2025

Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability

NeurIPS 2025poster

The development of effective explainability tools for Transformers is a crucial pursuit in deep learning research. One of the most promising approaches in this domain is Layer-wise Relevance Propagation (LRP), which propagates relevance scores backward through the network to the input space by redis…

Cited by 0SourceScholar
2025

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

ICML 2025oral

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence.…

Cited by 8SourcePDFScholar
2024

The Hidden Language of Diffusion Models

ICLR 2024poster

Text-to-image diffusion models have demonstrated an unparalleled ability to generate high-quality, diverse images from a textual prompt. However, the internal representations learned by these models remain an enigma. In this work, we present Conceptor, a novel method to interpret the internal repres…

2023

Discriminative Class Tokens for Text-to-Image Diffusion Models

ICCV 2023poster

Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in the input text. One way of alleviating these issues is to trai…

Cited by 10PDFcodeScholar
2022

No Token Left Behind: Explainability-Aided Image Classification and Generation

ECCV 2022poster

"The application of zero-shot learning in computer vision has been revolutionized by the use of image-text matching models. The most notable example, CLIP, has been widely used for both zero-shot classification and guiding generative models with a text prompt. However, the zero-shot use of CLIP is u…

2022

Optimizing Relevance Maps of Vision Transformers Improves Robustness

NeurIPS 2022accept

It has been observed that visual classification models often rely mostly on spurious cues such as the image background, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and direct the model to base its predictio…

2021

Generic Attention-Model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers

ICCV 2021poster

Transformers are increasingly dominating multi-modal reasoning tasks, such as visual question answering, achieving state-of-the-art results thanks to their ability to contextualize information using the self-attention and co-attention mechanisms. These attention modules also play a role in other com…

Cited by 392PDFcodeScholar