← Search

Ruilong Xing

2 accepted papers

2026

FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

ICLR 2026oral

Although Video Large Language Models (VLLMs) have shown remarkable capabilities in video understanding, they are required to process high volumes of visual tokens, causing significant computational inefficiency. Existing VLLMs acceleration frameworks usually compress spatial and temporal redundancy…

Cited by 0SourcecodeScholar
2025

C2AD: Dual Consistency Learning for Zero-Shot Anomaly Detection

ICASSP 2025accepted

Zero-shot anomaly detection (ZSAD) is dedicated to detecting anomalies without having any seen normal or abnormal samples for the target set. Existing approaches utilize the pre-trained CLIP to assess normality/abnormality by exploiting the similarity between images and text with the frozen visual e…

Cited by 0SourceScholar