← Search

K. R. Prajwal

6 accepted papers

2024

MusicFlow: Cascaded Flow Matching for Text Guided Music Generation

ICML 2024poster

We introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music audios, we construct two flow matching networks to model the conditional distribution of semantic and acoustic features. Ad…

Cited by 9SourcePDFScholar
2022

Automatic Dense Annotation of Large-Vocabulary Sign Language Videos

ECCV 2022poster

"Recently, sign language researchers have turned to sign language interpreted TV broadcasts, comprising (i) a video of continuous signing and (ii) subtitles corresponding to the audio content, as a readily available and large-scale source of training data. One key challenge in the usability of such…

Cited by 24SourcePDFScholar
2020

Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis

CVPR 2020poster

Humans involuntarily tend to infer parts of the conversation from lip movements when the speech is absent or corrupted by external noise. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate natural speech given only the lip movements of a speaker. Acknowledging t…

Cited by 130PDFcodeScholar