MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement
Generating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speech factors and the coarse granularity of existing control mechanisms. To overcome