FedCDWA: Decoupled Federated Prototype Distillation with Hierarchical Wasserstein Aggregation
Zhenshen Liu, Kai Fan, Wenjie Li, Kuan Zhang, HUI LI, Yintang Yang
Abstract
Federated learning enables decentralized clients to collaboratively train models without sharing local data. However, heterogeneous client distributions often induce client drift and hinder convergence. This paper proposes FedCDWA, a decoupled hierarchical federated distillation framework. FedCDWA decouples client-side personalized distillation from server-side mutual distillation to mitigate distillation-induced optimization conflicts. It further adopts Hierarchical Wasserstein Aggregation to aggregate prototypes without restrictive parametric assumptions while preserving intra-class structure and inter-class geometry. To achieve finer-grained feature alignment, Prototype–Variance Dual Alignment matches feature means and variances in the feature space. We prove convergence guarantees for FedCDWA. Experiments on three datasets demonstrate that FedCDWA consistently improves both global and personalized accuracy across heterogeneity levels, with smaller performance degradation under more severe heterogeneity.
BibTeX
@inproceedings{
liu2026fedcdwa,
title={Fed{CDWA}: Decoupled Federated Prototype Distillation with Hierarchical Wasserstein Aggregation},
author={Zhenshen Liu and Kai Fan and Wenjie Li and Kuan Zhang and HUI LI and Yintang Yang},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=k0qTjaVatI}
}