← Search

Jing-Jing Li

2 accepted papers

2026

PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

ICLR 2026poster

Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduc…

Cited by 0SourceScholar
2025

SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior

ICML 2025poster

The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community's values), which current systems fall short on. To address this gap, we present SafetyAnalyst, a novel AI sa…

Cited by 0SourcePDFScholar