← Search

Arjun Jagota

1 accepted papers

2026

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

ICLR 2026poster

The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge. While Inverse Reinforcement Learning (IRL) can infer reward functions from behaviour, existing approaches either produce a single, overconfiden…

Cited by 0SourceScholar