2026
Robust Preference Alignment via Directional Neighborhood Consensus
ICLR 2026poster
Aligning large language models with human preferences is critical for creating reliable and controllable AI systems. A human preference can be visualized as a high-dimensional vector where different directions represent trade-offs between desired attributes (e.g., helpfulness vs. verbosity). Yet, be…