← Search

Tom Everitt

13 accepted papers

2024

Discovering Agents (Abstract Reprint)

AAAI 2024technical

Causal models of agents have been used to analyse the safety aspects of machine learning systems. But identifying agents is non-trivial – often the causal model is just assumed by the modeller without much justification – and modelling failures can lead to mistakes in the safety analysis. This paper…

Cited by 0SourcePDFScholar
2024

Reasoning about Causality in Games (Abstract Reprint)

AAAI 2024technical

Causal reasoning and game-theoretic reasoning are fundamental topics in artificial intelligence, among many other disciplines: this paper is concerned with their intersection. Despite their importance, a formal framework that supports both these forms of reasoning has, until now, been lacking. We of…

Cited by 1SourcePDFScholar
2023

Honesty Is the Best Policy: Defining and Mitigating AI Deception

NeurIPS 2023spotlight

Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (for instance, in our experiments with language models, the goal of being evaluated as truthful). There are a number of e…

Cited by 35SourcePDFScholar
2022

A Complete Criterion for Value of Information in Soluble Influence Diagrams

AAAI 2022technical

Influence diagrams have recently been used to analyse the safety and fairness properties of AI systems. A key building block for this analysis is a graphical criterion for value of information (VoI). This paper establishes the first complete graphical criterion for VoI in influence diagrams with mul…

2022

Why Fair Labels Can Yield Unfair Predictions: Graphical Conditions for Introduced Unfairness

AAAI 2022technical

In addition to reproducing discriminatory relationships in the training data, machine learning (ML) systems can also introduce or amplify discriminatory effects. We refer to this as introduced unfairness, and investigate the conditions under which it may arise. To this end, we propose introduced tot…

2021

Agent Incentives: A Causal Perspective

AAAI 2021technical

We present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a new graphical criterion for value of control, establishing its soundness and completeness. We also introduce two new conc…