This work investigates two forms of compute in diffusion language models for text‑to‑speech (TTS): model depth (parameter count) and refinement steps $T$ (inference budget). We trained 15 masked‑diffusion codec TTS models with depths ranging from 19 M to 133 M parameters (three random seeds), using 2,000 h of speech data, and swept refinement steps $T \in [1,16]$ at inference time. Evaluation covered zero‑shot synthesis measured by ASR word error rate (intelligibility) and speaker verification accuracy (identity) on 174 held‑out speakers.
The results show that refinement closes 86.2% of the intelligibility error range but only 46.4% of the identity error range, yielding a 1.86× asymmetry that persists across multiple error metrics. Retraining with 3× and 6× schedules attenuates the gap (1.84 → 1.36 → 1.23) because intelligibility saturates with steps while identity continues to improve.
We also applied a Best‑of‑K search strategy: when refinement fails to recover speaker identity, sampling from four independent encoders yields win rates of 64.6%‑79.0%. Depth $d$ and steps $T$ are not interchangeable; a separable $B(d)B(T)$ model fits significantly better than substitution models (ΔAICc=+69.3), indicating they address different bottlenecks. Further analysis attributes 62% of the remaining identity deficit to the codec rather than the generator.
In summary, refinement primarily alleviates intelligibility bottlenecks, while depth mainly enhances identity preservation; they should be optimized separately rather than treated as substitutes.
Review: Refinement steps dramatically improve speech clarity but have limited impact on speaker identity, which can be rescued by search‑based sampling. Future efforts should allocate more compute to the codec to close the identity gap while maintaining intelligibility.