2026
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
ICLR 2026poster
The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge. While Inverse Reinforcement Learning (IRL) can infer reward functions from behaviour, existing approaches either produce a single, overconfiden…