ICASSP 2024accepted0 citations

GPT-4 Driven Cinematic Music Generation Through Text Processing

Muhammad Taimoor Haseeb, Ahmad Hammoudeh, Gus Xia

Abstract

This paper presents Herrmann-1 <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> , a multimodal framework to generate background music tailored to movie scenes, by integrating state-of-the-art vision, language, music, and speech processing models. Our pipeline begins by extracting visual and speech information from a movie scene, performing emotional analysis on it, and converting these into descriptive texts. Then, GPT-4 translates these high-level descriptions into low-level music conditions. Finally, these text-based music conditions guide a text-to-music model to generate music that resonates with input movie scenes. Comprehensive objective and subjective evaluations attest to the high synthesis quality, congruence, and superiority of our pipeline.

BibTeX
@inproceedings{icassp2024_gpt4drivencinema,
  title = {GPT-4 Driven Cinematic Music Generation Through Text Processing},
  author = {Muhammad Taimoor Haseeb and Ahmad Hammoudeh and Gus Xia},
  booktitle = {ICASSP 2024},
  year = {2024}
}