Hallucination-Aware Prompt Optimization for Text-to-Video Synthesis
Jiapeng Wang, Chengyu Wang, Jun Huang, Lianwen Jin
Abstract
The rapid advancements in AI-generated content (AIGC) have led to extensive research and application of deep text-to-video (T2V) synthesis models, such as OpenAI's Sora. These models typically rely on high-quality prompt-video pairs and detailed text prompts for model training in order to produce high-quality videos. To boost the effectiveness of Sora-like T2V models, we introduce VidPrompter, an innovative large multi-modal model supporting T2V applications with three key functionalities: (1) generating detailed prompts from raw videos, (2) enhancing prompts from videos grounded with short descriptions, and (3) refining simple user-provided prompts to elevate T2V video quality. We train VidPrompter using a hybrid multi-task paradigm and propose the hallucination-aware direct preference optimization (HDPO) technique to improve the multi-modal, multi-task prompt optimization process. Experiments on various tasks show our method surpasses strong baselines and other competitors.
BibTeX
@inproceedings{ijcai2025_hallucinationawa,
title = {Hallucination-Aware Prompt Optimization for Text-to-Video Synthesis},
author = {Jiapeng Wang and Chengyu Wang and Jun Huang and Lianwen Jin},
booktitle = {IJCAI 2025},
year = {2025}
}