← Search

Rami Ben-Ari

13 accepted papers

2026

DIFFUSION-BASED UNSUPERVISED AUDIO-VISUAL SPEECH SEPARATION IN NOISY ENVIRONMENTS WITH NOISE PRIOR

ICASSP 2026poster

In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models both clean speech and structured noise components, training exclusively on these individual signals rather than noisy mix…

Cited by 0SourcePDFScholar
2026

Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention

ICML 2026poster

Autoregressive video diffusion models enable \emph{streaming} generation, opening the door to long-form synthesis, video world models, and interactive neural game engines. However, their core attention layers become a major bottleneck at inference time: as generation progresses, the KV cache grows, …

Cited by 0SourceScholar
2025

CarGait: Cross-Attention based Re-ranking for Gait recognition

ICCV 2025poster

Gait recognition is a computer vision task that identifies individuals based on their walking patterns. Its performance is commonly evaluated by ranking a gallery of candidates and measuring the identification accuracy at Rank-K. Existing models are typically single-staged, searching for the probe's…

Cited by 0SourcePDFScholar
2025

EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition

ICLR 2025poster

The task of Visual Place Recognition (VPR) is to predict the location of a query image from a database of geo-tagged images. Recent studies in VPR have highlighted the significant advantage of employing pre-trained foundation models like DINOv2 for the VPR task. However, these models are often deeme…

Cited by 10SourcePDFScholar
2025

Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization

NeurIPS 2025poster

We address the challenge of Small Object Image Retrieval (SoIR), where the goal is to retrieve images containing a specific small object, in a cluttered scene. The key challenge in this setting is constructing a single image descriptor, for scalable and efficient search, that effectively represents…

Cited by 0SourceScholar
2025

Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models

ICLR 2025poster

Diffusion inversion is the problem of taking an image and a text prompt that describes it and finding a noise latent that would generate the exact same image. Most current deterministic inversion techniques operate by approximately solving an implicit equation and may converge slowly or yield poor…

2024

Data Roaming and Quality Assessment for Composed Image Retrieval

AAAI 2024technical

The task of Composed Image Retrieval (CoIR) involves queries that combine image and text modalities, allowing users to express their intent more effectively. However, current CoIR datasets are orders of magnitude smaller compared to other vision and language (V&L) datasets. Additionally, some of the…

2024

Generating Images of Rare Concepts Using Pre-trained Diffusion Models

AAAI 2024technical

Text-to-image diffusion models can synthesize high quality images, but they have various limitations. Here we highlight a common failure mode of these models, namely, generating uncommon concepts and structured concepts like hand palms. We show that their limitation is partly due to the long-tail na…

2024

Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval

NeurIPS 2024poster

Personalized retrieval and segmentation aim to locate specific instances within a dataset based on an input image and a short description of the reference instance. While supervised methods are effective, they require extensive labeled data for training. Recently, self-supervised foundation models h…

2023

Chatting Makes Perfect: Chat-based Image Retrieval

NeurIPS 2023poster

Chats emerge as an effective user-friendly approach for information retrieval, and are successfully employed in many domains, such as customer service, healthcare, and finance. However, existing image retrieval approaches typically address the case of a single query-to-image round, and the use of ch…

2023

Norm-guided latent space exploration for text-to-image generation

NeurIPS 2023poster

Text-to-image diffusion models show great potential in synthesizing a large variety of concepts in new compositions and scenarios. However, the latent space of initial seeds is still not well understood and its structure was shown to impact the generation of various concepts. Specifically, simple op…

2021

Noise Estimation Using Density Estimation for Self-Supervised Multimodal Learning

AAAI 2021technical

One of the key factors of enabling machine learning models to comprehend and solve real-world tasks is to leverage multimodal data. Unfortunately, annotation of multimodal data is challenging and expensive. Recently, self-supervised multimodal methods that combine vision and language were proposed t…