Beyond Client Clustering: Fine-Grained Preference Alignment in Federated RLHF via Self-Evolving Routing
Ke Wang, Shaojing Fu, Yuchuan Luo, Guilin Deng, Silong Chen, Zheng Yuan, Lin Liu
Abstract
Federated Reinforcement Learning from Human Feedback (RLHF) enables the collaborative alignment of Large Language Models (LLMs) while preserving privacy, yet it faces critical bottlenecks arising from data heterogeneity. Existing approaches typically rely on rigid client-level clustering, which overlooks intra-client heterogeneity and fails to adapt to the multifaceted needs of individual users. To address this, we propose FedPrism, a novel framework that shifts alignment granularity from the coarse client level to the precise instance level. Similar to an optical prism dispersing mixed light into distinct spectral components, FedPrism decomposes complex, intra-client heterogeneous data streams by dynamically routing individual samples to specialized experts based on semantic features. Crucially, we devise a Posterior-Guided Performance Distillation (PGPD) mechanism that leverages experts' actual training loss as self-supervision to autonomously refine routing policies without explicit labels. Extensive experiments show that FedPrism not only establishes new state-of-the-art results but also effectively mitigates the negative transfer prevalent in non-IID settings, ensuring superior alignment fidelity.
BibTeX
@inproceedings{ijcai2026_beyondclientclus,
title = {Beyond Client Clustering: Fine-Grained Preference Alignment in Federated RLHF via Self-Evolving Routing},
author = {Ke Wang and Shaojing Fu and Yuchuan Luo and Guilin Deng and Silong Chen and Zheng Yuan and Lin Liu},
booktitle = {IJCAI 2026},
year = {2026}
}