← Search

Yuval Kluger

22 accepted papers

2025

Dual Diffusion for Unified Image Generation and Understanding

CVPR 2025poster

Diffusion models have gained tremendous success in text-to-image generation, yet still struggle with visual understanding tasks, an area dominated by autoregressive vision-language models. We propose a large-scale and fully end-to-end diffusion model for multi-modal understanding and generation that…

Cited by 81SourcePDFScholar
2025

Partition First, Embed Later: Laplacian-Based Feature Partitioning for Refined Embedding and Visualization of High-Dimensional Data

ICML 2025oral

Embedding and visualization techniques are essential for analyzing high-dimensional data, but they often struggle with complex data governed by multiple latent variables, potentially distorting key structural characteristics. This paper considers scenarios where the observed features can be partitio…

Cited by 0SourcePDFScholar
2025

Understanding and Enhancing Mask-Based Pretraining towards Universal Representations

NeurIPS 2025poster

Mask-based pretraining has become a cornerstone of modern large-scale models across language, vision, and recently biology. Despite its empirical success, its role and limits in learning data representations have been unclear. In this work, we show that the behavior of mask-based pretraining can be…

Cited by 0SourcecodeScholar
2024

Hyperbolic Diffusion Procrustes Analysis for Intrinsic Representation of Hierarchical Data Sets

ICASSP 2024accepted

In this paper, we present Hyperbolic Diffusion Procrustes Analysis (HDPA), a new method for informative representation of hierarchical datasets based on hyperbolic geometry, diffusion geometry, and Procrustes analysis. Our method jointly embeds multiple datasets in a product manifold of hyperbolic s…

Cited by 0SourceScholar
2024

Likelihood Training of Cascaded Diffusion Models via Hierarchical Volume-preserving Maps

ICLR 2024spotlight

Cascaded models are multi-scale generative models with a marked capacity for producing perceptually impressive samples at high resolutions. In this work, we show that they can also be excellent likelihood models, so long as we overcome a fundamental difficulty with probabilistic multi-scale models:…

2024

Transductive and Inductive Outlier Detection with Robust Autoencoders

UAI 2024poster

Accurate detection of outliers is crucial for the success of numerous data analysis tasks. In this context, we propose the Probabilistic Robust AutoEncoder (PRAE) that can simultaneously remove outliers during training (transductive) and learn a mapping that can be used to detect outliers in new dat…

Cited by 2SourcePDFScholar
2023

Multi-modal differentiable unsupervised feature selection

UAI 2023poster

Multi-modal high throughput biological data presents a great scientific opportunity and a significant computational challenge. In multi-modal measurements, every sample is observed simultaneously by two or more sets of sensors. In such settings, many observed variables in both modalities are often n…

2022

Crowdsourcing Regression: A Spectral Approach

AISTATS 2022poster

Merging the predictions of multiple experts is a frequent task. When ground-truth response values are available, this merging is often based on the estimated accuracies of the experts. In various applications, however, the only available information are the experts’ predictions on unlabeled test dat…

Cited by 2SourcePDFScholar
2022

Locally Sparse Neural Networks for Tabular Biomedical Data

ICML 2022spotlight

Tabular datasets with low-sample-size or many variables are prevalent in biomedicine. Practitioners in this domain prefer linear or tree-based models over neural networks since the latter are harder to interpret and tend to overfit when applied to tabular datasets. To address these neural networks’…

2021

Differentiable Unsupervised Feature Selection based on a Gated Laplacian

NeurIPS 2021poster

Scientific observations may consist of a large number of variables (features). Selecting a subset of meaningful features is often crucial for identifying patterns hidden in the ambient space. In this paper, we present a method for unsupervised feature selection, and we demonstrate its advantage in c…

2018

Learning Binary Latent Variable Models: A Tensor Eigenpair Approach

ICML 2018oral

Latent variable models with hidden binary units appear in various applications. Learning such models, in particular in the presence of noise, is a challenging computational problem. In this paper we propose a novel spectral approach to this problem, based on the eigenvectors of both the second order…

2018

SpectralNet: Spectral Clustering using Deep Neural Networks

ICLR 2018poster

Spectral clustering is a leading and popular technique in unsupervised data analysis. Two of its major limitations are scalability and generalization of the spectral embedding (i.e., out-of-sample-extension). In this paper we introduce a deep learning approach to spectral clustering that overcomes…

2016

A Deep Learning Approach to Unsupervised Ensemble Learning

ICML 2016poster

We show how deep learning methods can be applied in the context of crowdsourcing and unsupervised ensemble learning. First, we prove that the popular model of Dawid and Skene, which assumes that all classifiers are conditionally independent, is \em equivalent to a Restricted Boltzmann Machine (RBM)…

2016

Unsupervised Ensemble Learning with Dependent Classifiers

AISTATS 2016poster

In unsupervised ensemble learning, one obtains predictions from multiple sources or classifiers, yet without knowing the reliability and expertise of each source, and with no labeled data to assess it. The task is to combine these possibly conflicting predictions into an accurate meta-learner. Most w…

Cited by 58SourcePDFScholar
2015

Estimating the accuracies of multiple classifiers without labeled data

AISTATS 2015poster

In various situations one is given only the predictions of multiple classifiers over a large unlabeled test data. This scenario raises the following questions: Without any labeled data and without any a-priori knowledge about the reliability of these different classifiers, is it possible to consiste…

Cited by 73SourcePDFScholar