Vision‑language models may rewrite anomalous text in images into linguistically plausible forms, compromising OCR transcription faithfulness. Sequence‑level task rewards and local teacher guidance complement each other, yet guidance from a fixed teacher becomes less effective as the student improves. Offline analysis shows that supervision from an immutable teacher degrades across training checkpoints and across response groups with higher task rewards. Motivated by this, we introduce GAD‑RL, which adaptively regulates teacher supervision during joint post‑training based on the student’s current task performance and local distributions. A frozen teacher conditions on reference transcriptions and student‑generated prefixes. GAD‑RL disables distillation for response groups whose output task reward reaches at least 0.95 and continuously attenuates distillation strength as the group‑mean reward rises. It also weights the forward KL by the student’s probability of the teacher’s top‑1 token, moderating local auxiliary updates when student support for that candidate is low. On Qwen3.5‑2B, GAD‑RL achieves 59.92% Micro Recall on CHAOS‑Bench, surpassing GRPO and GRPO+OPD (fixed‑weight) by 8.45 and 4.43 percentage points respectively, and attains an Overall score of 91.18 on OmniDocBench v1.6.
Review