← Search

Zhiding Xiao

1 accepted papers

2025

StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language Models

COLING 2025main

The rapid development of multimodal large language models (MLLMs) has positioned visual storytelling as a crucial area in content creation. However, existing models often struggle to maintain temporal, spatial, and narrative coherence across image sequences, and they frequently lack the depth and en…

Cited by 3SourcePDFScholar