← Search

Matthew Walmer

7 accepted papers

2026

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders

CVPR 2026

The space of task-agnostic feature upsampling has emerged as a promising area of research to efficiently create denser features from pre-trained visual backbones. These methods act as a shortcut to achieve dense features for a fraction of the cost by learning to map low-resolution features to high-r

Cited by 0SourcecodeScholar
2025

Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition

ICCV 2025poster

Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to…

Cited by 0SourcePDFScholar
2024

LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors

ECCV 2024poster

"We present a simple self-supervised method to enhance the performance of ViT features for dense downstream tasks. Our Lightweight Feature Transform (LiFT) is a straightforward and compact postprocessing network that can be applied to enhance the features of any pre-trained ViT backbone. LiFT is fas…

Cited by 5SourcePDFScholar
2023

TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models

ICCV 2023oral

We present a Multimodal Backdoor defense technique TIJO (Trigger Inversion using Joint Optimization). Recently Walmer et al. demonstrated successful backdoor attacks on multimodal models for the Visual Question Answering task. Their dual-key backdoor trigger is split across two modalities (image and…

Cited by 12PDFcodeScholar
2023

Teaching Matters: Investigating the Role of Supervision in Vision Transformers

CVPR 2023poster

Vision Transformers (ViTs) have gained significant popularity in recent years and have proliferated into many applications. However, their behavior under different learning paradigms is not well explored. We compare ViTs trained through different methods of supervision, and show that they learn a di…

2022

Dual-Key Multimodal Backdoors for Visual Question Answering

CVPR 2022poster

The success of deep learning has enabled advances in multimodal tasks that require non-trivial fusion of multiple input domains. Although multimodal models have shown potential in many problems, their increased complexity makes them more vulnerable to attacks. A Backdoor (or Trojan) attack is a clas…

Cited by 52PDFcodeScholar
2020

APRICOT: A Dataset of Physical Adversarial Attacks on Object Detection

ECCV 2020poster

Physical adversarial attacks threaten to fool object detection systems, but reproducible research on the real-world effectiveness of physical patches and how to defend against them requires a publicly available benchmark dataset. We present APRICOT, a collection of over 1,000 annotated photographs o…

Cited by 59SourcePDFScholar