← Search

Tom Bewley

9 accepted papers

2025

Interpreting Language Reward Models via Contrastive Explanations

ICLR 2025poster

Reward models (RMs) are a crucial component in the alignment of large language models’ (LLMs) outputs with human values. RMs approximate human preferences over possible LLM responses to the same prompt by predicting and comparing reward scores. However, as they are typically modified versions of LLM…

Cited by 0SourcePDFScholar
2025

Representation Consistency for Accurate and Coherent LLM Answer Aggregation

NeurIPS 2025poster

Test-time scaling improves large language models' (LLMs) performance by allocating more compute budget during inference. To achieve this, existing methods often require intricate modifications to prompting and sampling strategies. In this work, we introduce representation consistency (RC), a test-ti…

Cited by 0SourceScholar
2025

To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models

ICML 2025poster

We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually tuned steering strengths, often resulting in under or overstee…

Cited by 0SourcePDFScholar
2024

Counterfactual Metarules for Local and Global Recourse

ICML 2024poster

We introduce **T-CREx**, a novel model-agnostic method for local and global counterfactual explanation (CE), which summarises recourse options for both individuals and groups in the form of generalised rules. It leverages tree-based surrogate models to learn the counterfactual rules, alongside *meta…

Cited by 3SourcePDFScholar
2024

Sequential Harmful Shift Detection Without Labels

NeurIPS 2024poster

We introduce a novel approach for detecting distribution shifts that negatively impact the performance of machine learning models in continuous production environments, which requires no access to ground truth data labels. It builds upon the work of Podkopaev and Ramdas [2022], who address scenarios…

Cited by 1SourcePDFScholar
2022

Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance Learning

NeurIPS 2022accept

We generalise the problem of reward modelling (RM) for reinforcement learning (RL) to handle non-Markovian rewards. Existing work assumes that human evaluators observe each step in a trajectory independently when providing feedback on agent behaviour. In this work, we remove this assumption, extendi…

2021

TripleTree: A Versatile Interpretable Representation of Black Box Agents and their Environments

AAAI 2021technical

In explainable artificial intelligence, there is increasing interest in understanding the behaviour of autonomous agents to build trust and validate performance. Modern agent architectures, such as those trained by deep reinforcement learning, are currently so lacking in interpretable structure as t…

2019

On The Combination of Gamification and Crowd Computation in Industrial Automation and Robotics Applications

ICRA 2019poster

Autonomous intelligent systems outperform human workers in an expanding range of domains, typically those in which success is a function of speed, precision and repeatability. However, many cognitive tasks remain beyond the reach of automation. In this work, we propose the use of video games to crow…

Cited by 7SourceScholar