← Search

Adam Mahdi

7 accepted papers

2026

A Positive Case for Faithfulness: Explanations Help Predict Model Behavior

ICML 2026poster

LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, typically relying on identifying unfaithfulness via adversarial prompting or det…

Cited by 0SourceScholar
2026

It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

ICML 2026poster

Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their reliance on dynamic web content, however, makes them vulnerable to prompt injection attacks: adversarial instructions hidden in interface elements that persuad…

Cited by 0SourceScholar
2026

LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation

ICLR 2026poster

Frontier language models appear strong at solving reasoning problems, but their performance is often inflated by shortcuts such as memorisation and knowledge. We introduce LingOLY-TOO, a challenging reasoning benchmark of 6,995 questions that counters these shortcuts by applying expert-designed obfu…

Cited by 0SourcecodeScholar
2026

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

ICML 2026oral

Multimodal agents offer a compelling path to automating complex document-intensive workflows, yet a critical question remains: do these architectures demonstrate genuine strategic reasoning, or simply conduct stochastic trial-and-error search? To address this, we introduce Agentic Document VQA, a be…

Cited by 0SourceScholar
2025

How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis

EMNLP 2025

Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of algorithm, but prior explanations—attributing its effects solely to dampened toxic neurons in the MLP layers—are incomplete

Cited by 0SourcePDFScholar
2025

LLMs Don’t Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations

EMNLP 2025

To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated counterfactual explanations (SCEs), where a model explains its prediction by modifying the input such that it would have p

2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

NeurIPS 2025poster

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as `safety' and `robustness' requires strong construct validity, that is, having measures t…

Cited by 0SourceScholar