← Search

Yating Yu

3 accepted papers

2026

CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World

AAAI 2026technical

How far are deep models from real-world video anomaly understanding (VAU)? Current works typically emphasize detecting unexpected occurrences deviating from normal patterns or comprehending anomalous events with interpretable descriptions. However, they exhibit only a superficial comprehension of re

Cited by 0SourcePDFScholar
2025

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

AAAI 2025technical

Zero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding. However, finetuning CLIP directly for ZSAR yields suboptimal performance, given its inherent constraints in capturing essential temporal dynamics from both vision and text perspectives, especially wh…

2025

Learning to Generalize without Bias for Open-Vocabulary Action Recognition

ICCV 2025poster

Leveraging the effective visual-text alignment and static generalizability from CLIP, recent video learners adopt CLIP initialization with further regularization or recombination for generalization in open-vocabulary action recognition in-context. However, due to the static bias of CLIP, such video…