GEM: Generating Engaging Multimodal Content
Chongyang Gao, Yiren Jian, Natalia Denisenko, Soroush Vosoughi, V. S. Subrahmanian
Abstract
Generating engaging multimodal content is a key objective in numerous applications, such as the creation of online advertisements that captivate user attention through a synergy of images and text. In this paper, we introduce GEM, a novel framework engineered for the generation of engaging multimodal image-text posts. The GEM framework operates in two primary phases. Initially, GEM integrates a pre-trained engagement discriminator with a technique for deriving an effective continuous prompt tailored for the stable diffusion model. Subsequently, GEM unveils an iterative algorithm dedicated to producing coherent and compelling image-sentence pairs centered around a specified topic of interest. Through a combination of experimental analysis and human evaluations, we establish that the image-sentence pairs generated by GEM not only surpass several established baselines in terms of engagement but also in achieving superior alignment.
BibTeX
@inproceedings{ijcai2024p847,
title = {GEM: Generating Engaging Multimodal Content},
author = {Gao, Chongyang and Jian, Yiren and Denisenko, Natalia and Vosoughi, Soroush and Subrahmanian, V. S.},
booktitle = {Proceedings of the Thirty-Third International Joint Conference on
Artificial Intelligence, {IJCAI-24}},
publisher = {International Joint Conferences on Artificial Intelligence Organization},
editor = {Kate Larson},
pages = {7654--7662},
year = {2024},
month = {8},
note = {AI, Arts & Creativity},
doi = {10.24963/ijcai.2024/847},
url = {https://doi.org/10.24963/ijcai.2024/847},
}