AAAI 2026technical0 citations
VORTEX: Aligning Task Utility and Human Preferences Through LLM-Guided Reward Shaping
Abstract
In social impact optimization, AI decision systems often rely on solvers that optimize well-calibrated mathematical objectives. However, these solvers cannot directly accommodate evolving human preferences, typically expressed in natural language rather than formal constraints. Recent approaches address this by using large language models (LLMs) to generate new reward functions from preference descriptions. While flexible, they risk sacrificing the system
BibTeX
@inproceedings{aaai2026_vortexaligningta,
title = {VORTEX: Aligning Task Utility and Human Preferences Through LLM-Guided Reward Shaping},
author = {Guojun Xiong and Milind Tambe},
booktitle = {AAAI 2026},
year = {2026}
}