NeFut Logo NeFut
Admin Login

[CS.AI] Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

Published at: 2026-08-24 22:00 Last updated: 2026-08-29 12:04
#AI #Machine Learning #LLM

Specialist language models are typically evaluated by endpoint gains: a generalist scores lower, a specialist scores higher, and the gap is taken as evidence of specialization. This perspective overlooks the internal changes introduced by the model update itself. We propose a paired weight‑delta path audit and apply it to two public general‑to‑medical specialist checkpoint pairs: Gemma‑3‑4B‑IT → MedGemma‑4B‑IT and Qwen2.5‑7B‑Instruct → HuatuoGPT‑o1‑7B.

In both pairs, the full decoder‑side update strongly reconstructs the measured movement on medical benchmarks (endpoint‑normalized retention of 0.974 and 1.183, respectively), making each decoder delta a suitable substrate for the audit. However, the movement is not cleanly localized. The MLP family is the most prominent broad component in both pairs, yet mixed off‑domain shifts, 10‑seed matched controls, and endpoint‑anchored rollbacks prevent a unique coarse‑family explanation.

The audit therefore separates update‑level reconstruction from component‑level explanation. Its claims pertain only to text‑only multiple‑choice benchmark movement, not to clinical validation, repair, or circuit‑level mechanisms.

Blogger's Review: This work highlights that raw score improvements can mask intricate internal dynamics. By auditing weight deltas, we can pinpoint which sub‑networks truly drive medical specialization, offering a reproducible analysis framework even though a definitive component‑wise explanation remains elusive.

Original Source: https://arxiv.org/abs/2608.20768

[h] Back to Home