← Search

Ricardo Gutierrez-Osuna

10 accepted papers

2026

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

ICLR 2026poster

Real-time voice conversion and speaker anonymization require causal, low-latency synthesis without sacrificing intelligibility or naturalness. Current systems have a core representational mismatch: content is time-varying, while speaker identity is injected as a static global embedding. We introduce…

Cited by 0SourcecodeScholar
2022

Joint Hypoglycemia Prediction and Glucose Forecasting via Deep Multi-Task Learning

ICASSP 2022accepted

We present a multitask learning approach to the problem of hypoglycemia (HG) prediction in diabetes. The approach is based on a state-of-the-art time series forecasting model, N-BEATS, and extends it by adding a classification task so that the model performs both glucose forecasting (i.e., predictin…

Cited by 0SourceScholar
2022

Minimizing Residuals for Native-Nonnative Voice Conversion in a Sparse, Anchor-Based Representation of Speech

ICASSP 2022accepted

We present a dictionary-learning algorithm for reducing the sparse coding residual of an exemplar-based method for native-to-nonnative voice conversion (VC). The proposed algorithm iteratively updates the source and target speaker dictionaries to reduce both the residual and voice conversion error,…

Cited by 0SourceScholar
2021

A Sparse Coding Approach to Automatic Diet Monitoring with Continuous Glucose Monitors

ICASSP 2021accepted

Measuring dietary intake is a major challenge in the management of chronic diseases. Current methods rely on self-report measures, which are cumbersome to obtain and often unreliable. This article presents an approach to estimate dietary intake automatically by analyzing the post-prandial glucose re…

Cited by 0SourceScholar
2021

Towards The Development of Subject-Independent Inverse Metabolic Models

ICASSP 2021accepted

Diet monitoring is an important component of interventions in type 2 diabetes, but is time intensive and often inaccurate. To address this issue, we describe an approach to monitor diet automatically, by analyzing fluctuations in glucose after a meal is consumed. In particular, we evaluate three sta…

Cited by 0SourceScholar
2018

Accent Conversion Using Phonetic Posteriorgrams

ICASSP 2018accepted

Accent conversion (AC) aims to transform non-native speech to sound as if the speaker had a native accent. This can be achieved by mapping source spectra from a native speaker into the acoustic space of the non-native speaker. In prior work, we proposed an AC approach that matches frames between the…

Cited by 0SourceScholar
2018

Voice Conversion Through Residual Warping in a Sparse, Anchor-Based Representation of Speech

ICASSP 2018accepted

In previous work we presented a Sparse, Anchor-Based Representation of speech (SABR) that uses phonemic “anchors” to represent an utterance with a set of sparse non-negative weights. SABR is speaker-independent: combining weights from a source speaker with anchors from a target speaker can be used f…

Cited by 6SourceScholar
2016

Classification of bisyllabic lexical stress patterns in disordered speech using deep learning

ICASSP 2016accepted

Technology-based therapy tools can be of great benefit to children with developmental speech disabilities as they typically require sustained practice with a speech therapist for several years. Towards this aim, over the past 4 years we have developed speech processing tools to automatically detect…

Cited by 0SourceScholar
2015

Joint optimization of anatomical and gestural parameters in a physical vocal tract model

ICASSP 2015accepted

We describe a method for adapting a physical vocal tract model's anatomical and gestural parameters using acoustic information to match a target speaker. Physical vocal tract models are hard to adjust to match a speaker, as doing so requires information which is difficult to capture, such as X-Ray o…

Cited by 0SourceScholar