ECCV 2018poster52 citations

End-to-End Joint Semantic Segmentation of Actors and Actions in Video

Jingwei Ji, Shyamal Buch, Alvaro Soto, Juan Carlos Niebles

Abstract

Traditional video understanding tasks include human action recognition and actor/object semantic segmentation. However, the combined task of providing semantic segmentation for different actor classes simultaneously with their action class remains a challenging but necessary task for many applications. In this work, we propose a new end-to-end architecture for tackling this task in videos. Our model effectively leverages multiple input modalities, contextual information, and multitask learning in the video to directly output semantic segmentations in a single unified framework. We train and benchmark our model on the Actor-Action Dataset (A2D) for joint actor-action semantic segmentation, and demonstrate state-of-the-art performance for both segmentation and detection. We also perform experiments verifying our approach improves performance for zero-shot recognition, indicating generalizability of our jointly learned feature space.

BibTeX
@inproceedings{eccv2018_endtoendjointsem,
  title = {End-to-End Joint Semantic Segmentation of Actors and Actions in Video},
  author = {Jingwei Ji and Shyamal Buch and Alvaro Soto and Juan Carlos Niebles},
  booktitle = {ECCV 2018},
  year = {2018}
}
End-to-End Joint Semantic Segmentation of Actors and Actions in Video · ECCV 2018