IJCAI 2024poster3 citations

Retrieval Guided Music Captioning via Multimodal Prefixes

Nikita Srivatsan, Ke Chen, Shlomo Dubnov, Taylor Berg-Kirkpatrick

Abstract

In this paper we put forward a new approach to music captioning, the task of automatically generating natural language descriptions for songs. These descriptions are useful both for categorization and analysis, and also from an accessibility standpoint as they form an important component of closed captions for video content. Our method supplements an audio encoding with a retriever, allowing the decoder to condition on multimodal signal both from the audio of the song itself as well as a candidate caption identified by a nearest neighbor system. This lets us retain the advantages of a retrieval based approach while also allowing for the flexibility of a generative one. We evaluate this system on a dataset of 200k music-caption pairs scraped from Audiostock, a royalty-free music platform, and on MusicCaps, a dataset of 5.5k pairs. We demonstrate significant improvements over prior systems across both automatic metrics and human evaluation.

Application domains: Music and soundMethods and resources: Machine learning, deep learning, neural models, reinforcement learning
BibTeX
@inproceedings{ijcai2024p859,
  title     = {Retrieval Guided Music Captioning via Multimodal Prefixes},
  author    = {Srivatsan, Nikita and Chen, Ke and Dubnov, Shlomo and Berg-Kirkpatrick, Taylor},
  booktitle = {Proceedings of the Thirty-Third International Joint Conference on
               Artificial Intelligence, {IJCAI-24}},
  publisher = {International Joint Conferences on Artificial Intelligence Organization},
  editor    = {Kate Larson},
  pages     = {7762--7770},
  year      = {2024},
  month     = {8},
  note      = {AI, Arts & Creativity},
  doi       = {10.24963/ijcai.2024/859},
  url       = {https://doi.org/10.24963/ijcai.2024/859},
}
Retrieval Guided Music Captioning via Multimodal Prefixes · IJCAI 2024