We address adaptive computation in 3D medical image segmentation. Instead of engineering a stronger backbone, we let the optimizer decide how much spectral mixing each network stage requires. To this end we derive a two‑parameter operator family FHEAT from the discrete cosine transform (DCT) solution of a fractional heat equation. The operator is governed by a fractional order (\alpha) and a diffusion strength (D), and it reduces to the identity when (D=0).
By re‑parameterizing with the semigroup time (\tau = D\cdot\alpha), instances at the same resolution compose exactly. Consequently, any distribution of diffusion across same‑resolution stages is equivalent to a single Sobolev‑type regularizer whose strength is learned. This identity limit empowers each layer’s optimizer—not the designer—to decide whether global mixing is needed and how sharp it should be.
We instantiate FHEAT in a lightweight U‑shaped architecture (Light‑UNETR) coupled with a Kolmogorov‑Arnold mixer (KAN3D) and adaptive rational activations, yielding FHEAT‑Seg. Training on three public benchmarks with only 5%‑20% of labels triggers gradient‑driven spectral sparsification: seven of eight stage‑level operators drive (D) to zero, while the surviving decoder layer feeding the semi‑supervised attention map retains the sharpest low‑pass filter ((\alpha\approx0.9)). The retired layers become exact identity shortcuts at inference, cutting FLOPs from 4.29G to 0.90G (a 79% reduction) with only 0.975M parameters.
Under a standard semi‑supervised protocol, FHEAT‑Seg achieves Dice scores of 90.47% (left atrium), 78.79% (Pancreas‑CT), and 81.90% (BraTS 2019), outperforming five semi‑supervised methods and the Light‑UNETR baseline. Its larger variant also surpasses Light‑UNETR‑L under full supervision (Dice 93.09%, 85.11%, 87.19%) with 2.851M parameters and 55.75G FLOPs.
These results suggest that the allocation of spectral computation is a learnable property of optimization dynamics rather than a manually fixed design choice.
Review: The paper abstracts spectral mixing into a learnable regularizer via fractional diffusion, achieving notable FLOP savings and segmentation gains under scarce labeling, and opens a promising direction for adaptive medical image segmentation.