2024
Stepwise Alignment for Constrained Language Model Policy Optimization
NeurIPS 2024poster
Safety and trustworthiness are indispensable requirements for real-world applications of AI systems using large language models (LLMs). This paper formulates human value alignment as an optimization problem of the language model policy to maximize reward under a safety constraint, and then proposes…