← Search

Paul MONTAGUE

17 accepted papers

2026

Certified but Fooled! Breaking Certified Defenses with Ghost Certificates

AAAI 2026technical

Certified defenses promise provable robustness guarantees. We study the malicious exploitation of probabilistic certification frameworks to better understand the limits of guarantee provisions. Now, the objective is to not only mislead a classifier, but also to manipulate the certification process t

Cited by 0SourcePDFScholar
2026

Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement Learning

ICML 2026poster

Existing backdoor attacks on Reinforcement Learning (RL) typically rely on unrealistic white-box access to victim parameters, rewards, or observations. Inspired by real world behaviors, we introduce the Supply-Chain Backdoor (SCAB) attack to demonstrate that such assumptions are unnecessary. SCAB ta…

Cited by 0SourceScholar
2026

Semantic Robustness Certification for Vision-Language Models

ICML 2026poster

Vision-language models (VLMs) are now widely used in downstream tasks. However, real-world applications often expose VLMs to distribution shifts induced by semantic variation (e.g., shape, size, and style). Robustness certification determines if a model’s prediction changes when transformations are …

Cited by 0SourceScholar
2025

3D-Prover: Diversity Driven Theorem Proving With Determinantal Point Processes

NeurIPS 2025poster

A key challenge in automated formal reasoning is the intractable search space, which grows exponentially with the depth of the proof. This branching is caused by the large number of candidate proof tactics which can be applied to a given goal. Nonetheless, many of these tactics are semantically simi…

Cited by 0SourcecodeScholar
2025

Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them

ICLR 2025poster

Concept erasure has emerged as a promising technique for mitigating the risk of harmful content generation in diffusion models by selectively unlearning undesirable concepts. The common principle of previous works to remove a specific concept is to map it to a fixed generic concept, such as a neutra…

2025

Multi-level Certified Defense Against Poisoning Attacks in Offline Reinforcement Learning

ICLR 2025poster

Similar to other machine learning frameworks, Offline Reinforcement Learning (RL) is shown to be vulnerable to poisoning attacks, due to its reliance on externally sourced datasets, a vulnerability that is exacerbated by its sequential nature. To mitigate the risks posed by RL poisoning, we extend c…

Cited by 0SourcePDFScholar
2025

Position: Certified Robustness Does Not (Yet) Imply Model Security

ICML 2025oral

While certified robustness is widely promoted as a solution to adversarial examples in Artificial Intelligence systems, significant challenges remain before these techniques can be meaningfully deployed in real-world applications. We identify critical gaps in current research, including the paradox…

Cited by 0SourcePDFScholar
2024

BAIT: Benchmarking (Embedding) Architectures for Interactive Theorem-Proving

AAAI 2024technical

Artificial Intelligence for Theorem Proving (AITP) has given rise to a plethora of benchmarks and methodologies, particularly in Interactive Theorem Proving (ITP). Research in the area is fragmented, with a diverse set of approaches being spread across several ITP systems. This presents a significan…

2024

Erasing Undesirable Concepts in Diffusion Models with Adversarial Preservation

NeurIPS 2024poster

Diffusion models excel at generating visually striking content from text but can inadvertently produce undesirable or harmful content when trained on unfiltered internet data. A practical solution is to selectively removing target concepts from the model, but this may impact the remaining concepts.…

2024

Et Tu Certifications: Robustness Certificates Yield Better Adversarial Examples

ICML 2024poster

In guaranteeing the absence of adversarial examples in an instance's neighbourhood, certification mechanisms play an important role in demonstrating neural net robustness. In this paper, we ask if these certifications can compromise the very models they help to protect? Our new *Certification Aware…

2023

Enhancing the Antidote: Improved Pointwise Certifications against Poisoning Attacks

AAAI 2023technical

Poisoning attacks can disproportionately influence model behaviour by making small changes to the training corpus. While defences against specific poisoning attacks do exist, they in general do not provide any guarantees, leaving them potentially countered by novel attacks. In contrast, by examining…

Cited by 8SourcePDFScholar
2023

Feature-Space Bayesian Adversarial Learning Improved Malware Detector Robustness

AAAI 2023technical

We present a new algorithm to train a robust malware detector. Malware is a prolific problem and malware detectors are a front-line defense. Modern detectors rely on machine learning algorithms. Now, the adversarial objective is to devise alterations to the malware code to decrease the chance of bei…

Cited by 11SourcePDFScholar
2022

Double Bubble, Toil and Trouble: Enhancing Certified Robustness through Transitivity

NeurIPS 2022accept

In response to subtle adversarial examples flipping classifications of neural network models, recent research has promoted certified robustness as a solution. There, invariance of predictions to all norm-bounded attacks is achieved through randomised smoothing of network inputs. Today's state-of-the…

2022

On Global-view Based Defense via Adversarial Attack and Defense Risk Guaranteed Bounds

AISTATS 2022poster

It is well-known that deep neural networks (DNNs) are susceptible to adversarial attacks, which presents the most severe fragility of the deep learning system. Despite achieving impressive performance, most of the current state-of-the-art classifiers remain highly vulnerable to carefully crafted imp…

Cited by 8SourcePDFScholar
2021

Improving Ensemble Robustness by Collaboratively Promoting and Demoting Adversarial Robustness

AAAI 2021technical

Ensemble-based Adversarial Training is a principled approach to achieve robustness against adversarial attacks. An important technicality of this approach is to control the transferability of adversarial examples between ensemble members. We propose in this work a simple, but effective strategy to c…

2020

Improving Adversarial Robustness by Enforcing Local and Global Compactness

ECCV 2020poster

The fact that deep neural networks are susceptible to crafted perturbations severely impacts the use of deep learning in certain domains of application. Among many developed defense models against such attacks, adversarial training emerges as the most successful method that consistently resists a wi…

2019

Maximal Divergence Sequential Autoencoder for Binary Software Vulnerability Detection

ICLR 2019poster

Due to the sharp increase in the severity of the threat imposed by software vulnerabilities, the detection of vulnerabilities in binary code has become an important concern in the software industry, such as the embedded systems industry, and in the field of computer security. However, most of the wo…

Cited by 64SourcePDFScholar