← Search

Ziqian Ning

6 accepted papers

2025

Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation

AAAI 2025technical

Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core…

2025

StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

AAAI 2025technical

Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using language model-based or diffusion-based approaches, several challenges remain: 1) cur…

2024

Dualvc 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion

ICASSP 2024accepted

Voice conversion is becoming increasingly popular, and a growing number of application scenarios require models with streaming inference capabilities. The recently proposed DualVC attempts to achieve this objective through streaming model architecture design and intra-model knowledge distillation al…

Cited by 0SourceScholar
2024

Promptvc: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

ICASSP 2024accepted

Stylistic voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or reference speech to control the conversion process, which leads to limitations…

Cited by 0SourceScholar
2023

Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features

ICASSP 2023accepted

Voice conversion for highly expressive speech is challenging. Current approaches struggle with the balance between speaker similarity, intelligibility, and expressiveness. To address this problem, we propose Expressive-VC, a novel end-to-end voice conversion framework that leverages advantages from…

Cited by 0SourceScholar
2023

Preserving Background Sound in Noise-Robust Voice Conversion Via Multi-Task Learning

ICASSP 2023accepted

Background sound is an informative form of art that is helpful in providing a more immersive experience in real-application voice conversion (VC) scenarios. However, prior research about VC, mainly focusing on clean voices, pay rare attention to VC with background sound. The critical problem for pre…

Cited by 0SourceScholar