2025
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
ACL 2025long
When aligning large language models (LLMs), their performance across various tasks (such as being helpful, harmless, and honest) is heavily influenced by the composition of the training data. However, it is difficult to determine what mixture of data should be used to produce a model with strong per…