← Search

Li-Wei Chen

12 accepted papers

2025

A Variational Framework for Improving Naturalness in Generative Spoken Language Models

ICML 2025poster

The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized for autoregressive modeling. Speech tokens derived from self-supervised models (known as semantic tokens) typically focus…

2025

Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition

ICASSP 2025accepted

While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aw…

Cited by 0SourceScholar
2025

Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

ICASSP 2025accepted

Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmas…

Cited by 0SourceScholar
2025

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

ICASSP 2025accepted

Iterative self-training, or iterative pseudo-labeling (IPL)—using an improved model from the current iteration to provide pseudo-labels for the next iteration—has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker re…

Cited by 0SourceScholar
2023

A Unified One-Shot Prosody and Speaker Conversion System with Self-Supervised Discrete Speech Units

ICASSP 2023accepted

We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody and language content, leading to the degradation of naturalness in converted speech. Additionally, the lack of proper la…

Cited by 0SourceScholar
2023

A Vector Quantized Approach for Text to Speech Synthesis on Real-World Spontaneous Speech

AAAI 2023technical

Recent Text-to-Speech (TTS) systems trained on reading or acted corpora have achieved near human-level naturalness. The diversity of human speech, however, often goes beyond the coverage of these corpora. We believe the ability to handle such diversity is crucial for AI systems to achieve human-leve…

2023

Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings

ACL 2023short

The use of positional embeddings in transformer language models is widely accepted. However, recent research has called into question the necessity of such embeddings. We further extend this inquiry by demonstrating that a randomly initialized and frozen transformer language model, devoid of positio…

Cited by 14SourcePDFScholar
2023

Learning Similarity Metrics for Volumetric Simulations with Multiscale CNNs

AAAI 2023technical

Simulations that produce three-dimensional data are ubiquitous in science, ranging from fluid flows to plasma physics. We propose a similarity model based on entropy, which allows for the creation of physically meaningful ground truth distances for the similarity assessment of scalar and vectorial d…

2023

Speaker-Independent Acoustic-to-Articulatory Speech Inversion

ICASSP 2023accepted

To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising inversion target, since this space captures the mechanics of speech…

Cited by 0SourceScholar