2022
TP-VIT: A Two-Pathway Vision Transformer for Video Action Recognition
ICASSP 2022accepted
Recently, inspired by the success of Transformer in natural language processing tasks, a number of works have attempted to apply Transformer-based models to video action recognition. Existing works only use one RGB stream as the input for Transformer. How to use multiple pathways and multiple stream…