← Search

Richard P. Wildes

17 accepted papers

2024

Selective Interpretable and Motion Consistent Privacy Attribute Obfuscation for Action Recognition

CVPR 2024poster

Concerns for the privacy of individuals captured in public imagery have led to privacy-preserving action recognition. Existing approaches often suffer from issues arising through obfuscation being applied globally and a lack of interpretability. Global obfuscation hides privacy sensitive regions but…

Cited by 4SourcePDFScholar
2024

Visual Concept Connectome (VCC): Open World Concept Discovery and their Interlayer Connections in Deep Models

CVPR 2024highlight

Understanding what deep network models capture in their learned representations is a fundamental challenge in computer vision. We present a new methodology to understanding such vision models the Visual Concept Connectome (VCC) which discovers human interpretable concepts and their interlayer connec…

Cited by 8SourcePDFScholar
2023

MED-VT: Multiscale Encoder-Decoder Video Transformer With Application To Object Segmentation

CVPR 2023poster

Multiscale video transformers have been explored in a wide variety of vision tasks. To date, however, the multiscale processing has been confined to the encoder or decoder alone. We present a unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in videos. Multisca…

Cited by 22SourcePDFScholar
2023

StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos

CVPR 2023poster

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This motivates the need to temporally localize the instruction s…

Cited by 29SourcePDFScholar
2022

A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic Information

CVPR 2022poster

Deep spatiotemporal models are used in a variety of computer vision tasks, such as action recognition and video object segmentation. Currently, there is a limited understanding of what information is captured by these models in their intermediate representations. For example, while it has been obser…

Cited by 22PDFcodeScholar
2022

P3IV: Probabilistic Procedure Planning From Instructional Videos With Weak Supervision

CVPR 2022oral

In this paper, we study the problem of procedure planning in instructional videos. Here, an agent must produce a plausible sequence of actions that can transform the environment from a given start to a desired goal state. When learning procedure planning from instructional videos, most recent work l…

Cited by 52PDFcodeScholar
2018

A New Large Scale Dynamic Texture Dataset with Application to ConvNet Understanding

ECCV 2018poster

This paper introduces a new large scale dynamic texture dataset. The dataset is provided with two complementary organizations, one based on dynamics independent of spatial appearance and one based on spatial appearance independent of dynamics. With over 10,000 videos, the proposed Dynamic Texture Da…

Cited by 39SourcePDFScholar
2018

What Have We Learned From Deep Representations for Action Recognition?

CVPR 2018poster

As the success of deep models has led to their deployment in all areas of computer vision, it is increasingly important to understand how these representations work and what they are capturing. In this paper, we shed light on deep spatiotemporal representations by visualizing what two-stream mode…

Cited by 59SourcePDFScholar
2017

Spatiotemporal Multiplier Networks for Video Action Recognition

CVPR 2017poster

This paper presents a general ConvNet architecture for video action recognition based on multiplicative interactions of spacetime features. Our model combines the appearance and motion pathways of a two-stream architecture by motion gating and is trained end-to-end. We theoretically motivate multipl…

Cited by 1281PDFcodeScholar