2024
PeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action Recognition
CVPR 2024poster
Recent progress in Vision-Language (VL) foundation models has revealed the great advantages of cross-modality learning. However due to a large gap between vision and text they might not be able to sufficiently utilize the benefits of cross-modality information. In the field of human action recogniti…