← Search

Weikang Meng

3 accepted papers

2026

Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining

CVPR 2026

Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer from two fundamental limitations: severe visual information lo

Cited by 0SourcecodeScholar
2026

Norm$\times$Direction: Restoring the Missing Query Norm in Vision Linear Attention

ICML 2026poster

Linear attention mitigates the quadratic complexity of softmax attention but suffers from a critical loss of expressiveness. We identify two primary causes: (1) The normalization operation cancels the query norm, which breaks the correlation between a query's norm and the spikiness (entropy) of the …

Cited by 0SourceScholar
2025

PolaFormer: Polarity-aware Linear Attention for Vision Transformers

ICLR 2025poster

Linear attention has emerged as a promising alternative to softmax-based attention, leveraging kernelized feature maps to reduce complexity from quadratic to linear in sequence length. However, the non-negative constraint on feature maps and the relaxed exponential function used in approximation lea…

Cited by 2SourcePDFScholar