← Search

Linqin Wang

5 accepted papers

2025

SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models

AAAI 2025technical

With the rapid advancement of large language models (LLMs), discrete speech representations have become crucial for integrating speech into LLMs. Existing methods for speech representation discretization rely on a predefined codebook size and Euclidean distance-based quantization. However, 1) the si…

2025

Universal Low-Resource Speech Synthesis Via Phoneme Fusion Coordinating Low-Rank Decomposition

ICASSP 2025accepted

Recent advancements in end-to-end text-to-speech models have made significant progress. However, these approaches based on high-resource languages, are inapplicable for low-resource languages, and existing low-resource speech synthesis methods are typically specific to single languages. Consequently…

Cited by 0SourceScholar
2025

Voice Conversion via Structural Entropy

ICASSP 2025accepted

Voice conversion (VC) aims to transform a person’s voice to resemble that of another person while maintaining the original linguistic content. Existing methods suffer from the blurring of speech representations and the leakage of prosody information. To address this issue, this study introduces SEVC…

Cited by 0SourceScholar
2024

DETS: End-to-End Single-Stage Text-to-Speech Via Hierarchical Diffusion Gan Models

ICASSP 2024accepted

End-to-end single-stage text-to-speech models have garnered significant attention in recent research, surpassing the performance of conventional two-stage pipeline systems. While prior single-stage models have made substantial advancements, there remains room for improvement in addressing intermitte…

Cited by 0SourceScholar
2023

Non-parallel Accent Transfer based on Fine-grained Controllable Accent Modelling

EMNLP 2023long findings

Existing accent transfer works rely on parallel data or speech recognition models. This paper focuses on the practical application of accent transfer and aims to implement accent transfer using non-parallel datasets. The study has encountered the challenge of speech representation disentanglement an…

Cited by 0SourceScholar