In decentralized, server‑less learning scenarios devices often run heterogeneous model architectures, making the standard decentralized SGD undefined because it requires averaging parameters of different dimensions. Knowledge distillation (KD) sidesteps this issue by exchanging soft predictions instead of weights, but a convergence theory for fully decentralized, asynchronous peer‑to‑peer (P2P) KD has been missing. This work shifts consensus from the parameter space to the function (output) space: a KD interaction acts as a geometric contraction operator on peers’ predictive distributions in logit space, which we analyze in the Hilbert space of predictions w.r.t. a reference measure. Under usual smoothness and variance assumptions together with two realizability conditions—one linking parameter‑SGD to the functional step and another bounding the restricted task/KD alignment—the time‑averaged functional stationarity and function‑space disagreement converge at rate $O\bigl(1/(\eta T)\bigr)$ to a neighbourhood of size $O(\eta)+O(B_f^2)+O(\zeta_f^2)$. Here $B_f$ is the distance from the task optimum to the reachable model classes of the peers, and $\zeta_f$ measures persistent local‑task heterogeneity. Experiments on homogeneous, width‑heterogeneous, and mixed‑family networks show that KD contracts function disagreement by $40\sim61\times$, while isolated training does not. The sampled stationarity diagnostic exhibits late‑stage transient exponents between $0.99$ and $1.90$ on the shared‑skeleton main runs, and a four‑point stepsize sweep reveals the predicted transient: a trade‑off in neighbourhood size.
Review