LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
Xiaodong Wang, Langling Huang, Zhirong Wu, Xu Zhao, Teng Xu, Xuhong Xia, Peixi Peng
Abstract
The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings. To fill this gap, this paper proposes the first omnimodal benchmark for interactive livestream videos, LiViBench. It features a diverse set of 24 tasks, highlighting the perceptual, reasoning, and livestream-specific challenges. To efficiently construct the dataset, we design a standardized semi-automatic annotation workflow that incorporates the human-in-the-loop at multiple stages. The workflow leverages multiple MLLMs to form a multi-agent system for comprehensive video description and uses a seed-question-driven method to construct high-quality annotations. All interactive videos in the benchmark include audio, speech, and real-time comments modalities. To enhance models
BibTeX
@inproceedings{aaai2026_livibenchanomnim,
title = {LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding},
author = {Xiaodong Wang and Langling Huang and Zhirong Wu and Xu Zhao and Teng Xu and Xuhong Xia and Peixi Peng},
booktitle = {AAAI 2026},
year = {2026}
}