← Search

Adam Finkelstein

14 accepted papers

2024

Corn: Co-Trained Full- and No-Reference Speech Quality Assessment

ICASSP 2024accepted

Perceptual evaluation constitutes a crucial aspect of various audio-processing tasks. Full reference (FR) or similarity-based metrics rely on high-quality reference recordings, to which lower-quality or corrupted versions of the recording may be compared for evaluation. In contrast, no-reference (NR…

Cited by 0SourceScholar
2024

GR0: Self-Supervised Global Representation Learning for Zero-Shot Voice Conversion

ICASSP 2024accepted

Research in generative self-supervised learning (SSL) has largely focused on local embeddings for tokenized sequences. We introduce a generative SSL framework that learns a global representation that is disentangled from local embeddings. We apply this technique to jointly learn a global speaker emb…

Cited by 0SourceScholar
2022

Controllable Speech Representation Learning Via Voice Conversion and AIC Loss

ICASSP 2022accepted

Speech representation learning transforms speech into features that are suitable for downstream tasks, e.g. speech recognition, phoneme classification, or speaker identification. For such recognition tasks, a representation can be lossy (non-invertible), which is typical of BERT-like self-supervised…

Cited by 0SourceScholar
2021

CDPAM: Contrastive Learning for Perceptual Audio Similarity

ICASSP 2021accepted

Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function. The DPAM approach of Manocha et al. [1] learns a full-reference metric trained directly on human judgments, and thus correlates well with human perception. However, it re…

Cited by 0SourceScholar
2018

PairedCycleGAN: Asymmetric Style Transfer for Applying and Removing Makeup

CVPR 2018poster

This paper introduces an automatic method for editing a portrait photo so that the subject appears to be wearing makeup in the style of another person in a reference photo. Our unsupervised learning approach relies on a new framework of cycle-consistent generative adversarial networks. Different fro…

Cited by 348SourcePDFScholar
2016

Cute: A concatenative method for voice conversion using exemplar-based unit selection

ICASSP 2016accepted

State-of-the art voice conversion methods re-synthesize voice from spectral representations such as MFCCs and STRAIGHT, thereby introducing muffled artifacts. We propose a method that circumvents this concern using concatenative synthesis coupled with exemplar-based unit selection. Given parallel sp…

Cited by 0SourceScholar