← Search

Zixuan Xiong

3 accepted papers

2025

Efficient Visual Storytelling through Descriptive Words Distillation and Dynamic Decoding

ICASSP 2025accepted

Visual storytelling, a complex task in natural language generation, aims to create coherent and engaging narratives from a sequence of images, requiring more intricate and lengthy descriptions than typical image captioning. Current methods generally employ sophisticated modal interaction modules and…

Cited by 0SourceScholar
2025

Frozen Language Models Are Gradient Coherence Rectifiers in Vision Transformers

AAAI 2025technical

Large language models (LLMs) have demonstrated remarkable performance in multimodal tasks even with frozen LLM Block and only a few trainable parameters. However, the underlying mechanisms of how LLMs enhance multimodal performance remains unclear. In this work, we focus on the phenomenon that ``Mer…

Cited by 0SourcePDFScholar
2025

Youku Dense Caption: A Large-scale Chinese Video Dense Caption Dataset and Benchmarks

ICLR 2025poster

With the explosive growth of video content, video captions have emerged as a crucial tool for video comprehension, significantly enhancing the ability to understand and retrieve information from videos. However, most publicly available dense video captioning datasets are in English, resulting in a s…

Cited by 0SourcePDFScholar