2026
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
ICLR 2026poster
Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduc…