2025
StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language Models
COLING 2025main
The rapid development of multimodal large language models (MLLMs) has positioned visual storytelling as a crucial area in content creation. However, existing models often struggle to maintain temporal, spatial, and narrative coherence across image sequences, and they frequently lack the depth and en…