2022
MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration
ECCV 2022poster
"Multimodal video-audio-text understanding and generation can benefit from datasets that are narrow but rich. The narrowness allows bite-sized challenges that the research community can make progress on. The richness ensures we are making progress along the core challenges. To this end, we present a…