← Search

Longteng Kong

3 accepted papers

2026

CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics Priority

CVPR 2026

Cross-modal alignment aims to learn semantically consistent latent representations across diverse modalities. Prevailing methods rely on a text-guided aggregation paradigm to achieve fine-grained alignment, while they suffer from redundant patch-word correlations and high computational costs. To add

Cited by 0SourceScholar
2018

Hierarchical Attention and Context Modeling for Group Activity Recognition

ICASSP 2018accepted

Group activity recognition in videos is a challenging task, with two major issues, i.e. attending to those persons and their body parts that contribute significantly to the activity, and modeling contextual person structures in the group. Most previous approaches fail to provide a practical solution…

Cited by 0SourceScholar