Measuring the accent difference between two speakers is a fundamental problem in linguistics and speech technology. The chosen measurement depends on the research focus. A phonetics researcher may illustrate accent variation by comparing vowel formants in paired word recordings; this approach is interpretable but costly and may not reflect connected speech. In contrast, accent‑TTS work often relies on accent embeddings derived from classification tasks, which can be extracted from any recording yet remain opaque. This paper shows that articulatory representations obtained via articulatory inversion serve as an interpretable basis for accent comparison, and that optimal transport offers a unified framework to compare accents across arbitrary recording types. The approach combines flexibility with interpretability, providing a common tool for accent analysis and synthesis.
Review