2026
The Realignment Problem: When Right becomes Wrong in LLMs
ICML 2026poster
Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over time. Cultural shifts, value reinterpretations, and regulatory or industrial updates make static alignment increasingly brittle. As policies evolve, deploy…