← Search

Xiang Lyu

3 accepted papers

2026

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

ICLR 2026poster

Recent studies on end-to-end (E2E) speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs to generate discrete speech tokens. Existing E2E approaches primarily fall into two categories: (1) Methods that genera…

Cited by 0SourceScholar
2025

Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice

ICASSP 2025accepted

LLM-based text-to-speech(TTS) system has becoming the new trend and SOTA due to its high naturalness and zero-shot capability. However, it relies heavily on training data, usually requires at least thousands hours of labeled audio. In this report, we describe how to use pretrained CosyVoice model, t…

Cited by 0SourceScholar
2025

Fast Adaptation of Pretrained Speaker Verification System for Source Speaker Tracking

ICASSP 2025accepted

Traditional speaker verification system aims at distinguish speaker identity in real world audio, and has achieved satisfying performance in many scenarios. However, it is also very vulnerable, and can be easily attacked by voice anonymization system. In this report, we describe how to fast adapt a…

Cited by 0SourceScholar