2020
From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech
ICLR 2020poster
This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations. To this end, we propose a multi-modal learning framework that links the inference stage and generation stage. First, the inference networks are trained…