2025
Sparta Alignment: Collectively Aligning Multiple Language Models through Combat
NeurIPS 2025poster
We propose Sparta Alignment, an algorithm to collectively align multiple LLMs through competition and combat. To complement a single model's lack of diversity in generation and biases in evaluation, multiple LLMs form a 'sparta tribe' to compete against each other in fulfilling instructions while se…