Relexicalization is a pivotal technique in clinical NLP, enabling robust masking of sensitive information while generating datasets that retain real‑world characteristics. Preserving structural integrity, relational coherence, and temporal consistency during transformation remains challenging. Existing methods often replace entities independently, causing clinical inconsistencies across longitudinal records and reducing the utility of relexicalized data for downstream analysis.
To overcome these limitations, we introduce G-RELIC (Graph‑Based Contextual Relexicalization with Improved Consistency), which combines the generative power of large language models with graph structures. G-RELIC employs a graph‑based mapping mechanism to enforce a one‑to‑one correspondence between original and surrogate entities, and adds a deterministic temporal repositioning algorithm to maintain chronological consistency.
Empirical evaluations on diverse real‑world clinical datasets show that G-RELIC improves relational integrity by 30.4 percentage points (62.1% → 92.5%) and temporal coherence by 45.9 percentage points (46% → 91.9%), without violating established privacy benchmarks. This maximizes the analytical utility of relexicalized datasets while minimizing re‑identification risk.
Review: G-RELIC’s graph‑constrained and deterministic time handling delivers high‑consistency clinical document relexicalization, providing a more reliable foundation for subsequent research.