← Search

Yochai Yemini

4 accepted papers

2026

DIFFUSION-BASED UNSUPERVISED AUDIO-VISUAL SPEECH SEPARATION IN NOISY ENVIRONMENTS WITH NOISE PRIOR

ICASSP 2026poster

In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models both clean speech and structured noise components, training exclusively on these individual signals rather than noisy mix…

Cited by 0SourcePDFScholar
2024

LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading

ICLR 2024poster

Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for challenging and realistic datasets such as LRS3. In this work…

2021

GP-Tree: A Gaussian Process Classifier for Few-Shot Incremental Learning

ICML 2021spotlight

Gaussian processes (GPs) are non-parametric, flexible, models that work well in many tasks. Combining GPs with deep learning methods via deep kernel learning (DKL) is especially compelling due to the strong representational power induced by the network. However, inference in GPs, whether with or wit…

2020

A Composite DNN Architecture for Speech Enhancement

ICASSP 2020accepted

In speech enhancement, the use of supervised algorithms in the form of deep neural networks (DNNs) has become tremendously popular in recent years. The target function of the DNN (and the associated estimators) is often either a masking function applied to the noisy spectrum, or the clean log-spectr…

Cited by 0SourceScholar