Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods—based on saliency—produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Langu…