AAAI 2026technical0 citations

VORTEX: Aligning Task Utility and Human Preferences Through LLM-Guided Reward Shaping

Guojun Xiong, Milind Tambe

Abstract

In social impact optimization, AI decision systems often rely on solvers that optimize well-calibrated mathematical objectives. However, these solvers cannot directly accommodate evolving human preferences, typically expressed in natural language rather than formal constraints. Recent approaches address this by using large language models (LLMs) to generate new reward functions from preference descriptions. While flexible, they risk sacrificing the system

BibTeX
@inproceedings{aaai2026_vortexaligningta,
  title = {VORTEX: Aligning Task Utility and Human Preferences Through LLM-Guided Reward Shaping},
  author = {Guojun Xiong and Milind Tambe},
  booktitle = {AAAI 2026},
  year = {2026}
}