← Search

Brian Hu

3 accepted papers

2026

Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping

AAAI 2026technical

The deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments. Agents trained solely to achieve their objectives may adopt harmful behavior, exposing a key trade-off between maximizi

Cited by 0SourcePDFScholar
2024

Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain

NAACL 2024industry

In difficult decision-making scenarios, it is common to have conflicting opinions among expert human decision-makers as there may not be a single right answer. Such decisions may be guided by different attributes that can be used to characterize an individual’s decision. We introduce a novel dataset…

2023

Xaitk-Saliency: An Open Source Explainable AI Toolkit for Saliency

AAAI 2023technical

Advances in artificial intelligence (AI) using techniques such as deep learning have fueled the recent progress in fields such as computer vision. However, these algorithms are still often viewed as "black boxes", which cannot easily explain how they arrived at their final output decisions. Saliency…