← Search

Langling Huang

3 accepted papers

2026

A Training-Free Framework for Long Video Understanding via Video-Query-Options Similarity

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in image and short video understanding tasks, but their performance on hour-long videos remains limited due to constraint of input token capacity. Existing approaches often require costly training procedures, hindering their a…

Cited by 0SourceScholar
2026

Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning

CVPR 2026

Multimodal Large Language Models (MLLMs) have made great progress in video understanding tasks. However, when it comes to understanding complex or lengthy videos, MLLMs tend to overlook details or produce hallucinations. To alleviate these issues, recent work has attempted to leverage reinforcement

Cited by 0SourcecodeScholar
2026

LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding

AAAI 2026technical

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings. To fill this gap, this paper proposes the first omnimodal benchmark for intera

Cited by 0SourcePDFScholar