← Search

Yunyun Wang

7 accepted papers

2026

The Finer the Better: Towards Granular-aware Open-set Domain Generalization

AAAI 2026technical

Open-Set Domain Generalization (OSDG) aims to generalize over unseen target domains containing open classes, and the core challenge lies in identifying unknown samples never encountered during training. Recently, CLIP has exhibited impressive performance in OSDG, while it still falls into the dilemm

Cited by 0SourcePDFScholar
2025

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference

ICASSP 2025accepted

We propose a unified framework for Singing Voice Synthesis (SVS) and Conversion (SVC), addressing the limitations of existing approaches in cross-domain SVS/SVC, poor output musicality, and scarcity of singing data. Our framework enables control over multiple aspects, including language content base…

Cited by 0SourceScholar
2025

ProMEA: Prompt-driven Expansion and Alignment for Single Domain Generalization

IJCAI 2025

In single Domain Generalization (single-DG), data scarcity in the single source domain hampers the learning for invariant features, leading to overfitting over source domain and poor generalization to unseen target domains. Existing single-DG methods primarily augment the source domain by adversaria

Cited by 0SourcePDFScholar
2024

GR0: Self-Supervised Global Representation Learning for Zero-Shot Voice Conversion

ICASSP 2024accepted

Research in generative self-supervised learning (SSL) has largely focused on local embeddings for tokenized sequences. We introduce a generative SSL framework that learns a global representation that is disentangled from local embeddings. We apply this technique to jointly learn a global speaker emb…

Cited by 0SourceScholar
2022

Controllable Speech Representation Learning Via Voice Conversion and AIC Loss

ICASSP 2022accepted

Speech representation learning transforms speech into features that are suitable for downstream tasks, e.g. speech recognition, phoneme classification, or speaker identification. For such recognition tasks, a representation can be lossy (non-invertible), which is typical of BERT-like self-supervised…

Cited by 0SourceScholar
2020

Deep Audio Priors Emerge From Harmonic Convolutional Networks

ICLR 2020poster

Convolutional neural networks (CNNs) excel in image recognition and generation. Among many efforts to explain their effectiveness, experiments show that CNNs carry strong inductive biases that capture natural image priors. Do deep networks also have inductive biases for audio signals? In this paper,…

Cited by 40SourceScholar