Most Vision‑Language Models (VLMs) are built by extending a pretrained Large Language Model (LLM) with visual modules and multimodal alignment. This scaling often weakens the language‑side reasoning that the base LLM originally encodes. Although the base LLM retains usable reasoning after scaling, the aligned VLM cannot reliably access it. To address this, we propose LIFT (Language‑side reasonIng Facilitation and Transfer), a lightweight vector‑intervention technique that transfers reasoning capability from the base LLM to the VLM without retraining the backbone. LIFT defines a Reasoning Vector as the hidden‑state difference of the answer token between a Reasoner path that follows an explicit reasoning trace and a Solver path that skips it, and injects this vector into the language‑side activations of the target VLM. The method also allows learnable adaptation of the vector while keeping the VLM backbone frozen. We evaluate LIFT on two VLMs across six reasoning benchmarks, comparing vectors extracted from the base LLM and from the aligned VLM under matched protocols. Results consistently show that LLM‑derived vectors outperform VLM‑derived ones, confirming the base LLM as a more effective source for recovering reasoning. LIFT partially restores degraded reasoning through lightweight language‑side interventions, and analysis indicates that the vectors influence intermediate reasoning behavior rather than merely changing final answers. The source code will be released soon.
Review