Multimodal federated learning aims to leverage distributed data to improve model performance. Existing approaches assume homogeneous models and balanced modality distributions, which limits their applicability to real‑world settings with heterogeneous client architectures and severe modality imbalance. To tackle these issues we introduce the Multimodal Federated Learning Prototype‑guided Bilateral Alignment (MFedPBA) framework. MFedPBA achieves robust knowledge synergy through a dual alignment strategy.
At the feature level, a projection encoder maps heterogeneous feature spaces into a common representation, optimized by contrastive learning together with the Gromov‑Wasserstein distance.
At the decision level, naturally aligned logit prototypes are aggregated with entropy‑weighted averaging. This design jointly addresses heterogeneous feature spaces and collective decision aggregation.
Extensive experiments demonstrate that under model heterogeneity and modality imbalance, MFedPBA significantly outperforms state‑of‑the‑art baselines. Review