Text‑to‑speech systems are increasingly required to handle user‑generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical form rather than the surface spelling. We introduce UGTPhon, the first grapheme‑to‑phoneme (G2P) benchmark covering English, Vietnamese, and Korean, together with an inference‑grounded taxonomy for fine‑grained error diagnosis. Experiments reveal that existing G2P models and state‑of‑the‑art large language models (LLMs) suffer a systematic canonical‑to‑non‑canonical performance gap, reaching up to 66.8 PER (phoneme error rate) points. As a baseline, we propose a simple compositional G2P approach that first retrieves canonical‑form evidence via exact‑match lookup and then performs staged decoding. Across both ByT5 and Qwen2.5‑0.5B backbones, explicit modeling of the canonical form consistently reduces non‑canonical G2P errors. Notably, the 0.5B variant remains competitive with much larger few‑shot frontier LLMs, highlighting the benefit of explicitly modeling canonical‑form inference for UGT phonemization.
Review: This work provides a much‑needed multilingual benchmark and a detailed taxonomy that expose the shortcomings of current models on non‑canonical inputs, while the lightweight compositional method delivers substantial gains, offering a clear roadmap for future research.