ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understanding
We revisit video hallucination in multimodal large language models (Video-MLLMs) from a semantic aggregation perspective. While prior work attributes hallucinations to language priors, missing frames, or visual encoder biases, these explanations overlook errors arising during the aggregation of corr