← Search

Jana Schaich Borg

6 accepted papers

2026

Moral Change or Noise? On Problems of Aligning AI with Temporally Unstable Human Feedback

AAAI 2026technical

Alignment methods in moral domains seek to elicit moral preferences of human stakeholders and incorporate them into AI. This presupposes moral preferences as static targets, but such preferences often evolve over time. Proper alignment of AI to dynamic human preferences should ideally account for "l

Cited by 0SourcePDFScholar
2026

Position: We Need Practical AI Alignment Methods that Mirror Human Reasoning

ICML 2026poster

AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully c…

Cited by 0SourceScholar
2026

Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment

ICLR 2026poster

Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using standard preference elicitation methods, researchers and practitioners build models of human decisions and judgments, to which AI models are aligned. However…

Cited by 0SourceScholar
2025

SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior

ICML 2025poster

The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community's values), which current systems fall short on. To address this gap, we present SafetyAnalyst, a novel AI sa…

Cited by 0SourcePDFScholar
2025

Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics

EMNLP 2025

As large language models (LLMs) are increasingly used in morally sensitive domains, it is crucial to understand how persona traits affect their moral reasoning and persuasive behavior. We present the first large-scale study of multi-dimensional persona effects in AI-AI debates over real-world moral

Cited by 0SourcePDFScholar
2021

Indecision Modeling

AAAI 2021technical

AI systems are often used to make or contribute to important decisions in a growing range of applications, including criminal justice, hiring, and medicine. Since these decisions impact human lives, it is important that the AI systems act in ways which align with human values. Techniques for prefere…