NeFut Logo NeFut
中 Admin Login

[CS.AI] JIVEAdapter: A Multi-Task Additive Low-Rank Adapter via Joint and Individual Variation Explained

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Parameter‑efficient fine‑tuning (PEFT) updates only a tiny fraction of a pretrained model’s weights, dramatically reducing the cost compared with full fine‑tuning. Most existing low‑rank adapters are single‑task and apply multiplicative updates, so they do not explicitly model what information is shared across tasks and what is task‑specific.\ \ JIVEAdapter draws inspiration from the statistical method Joint and Individual Variation Explained (JIVE) and proposes an additive multi‑task low‑rank adapter. Each weight update @@@MATH_BLOCK0@@@ is decomposed as\ \ $$ \Delta W = \underbrace{J}{\text{Joint structure}} + \underbrace{I^{(t)}}_{\text{Individual structure for task } t} $$\ \ where $J$ is a shared Joint component across all tasks and $I^{(t)}$ is the Individual component for task $t$. To keep the two components interpretable and separate, JIVEAdapter penalises the Individual part to be nearly orthogonal to the Joint part during training:\ \ $$ \langle J, I^{(t)} \rangle \approx 0 $$\ \ The adapter’s rank is allocated adaptively between a shared Joint pool and per‑task Individual pools, allowing the most economical representation of both shared and task‑specific signals.\ \ The Joint component can be learned either jointly over a whole task group or incrementally: first learn and freeze $J$ on existing tasks, then for a new task only learn $I^{(t)}$ and a per‑direction scale $\alpha^{(t)}$, without retraining the shared part.\ \ Experiments on GLUE and SuperGLUE using DeBERTaV3‑base show that, at the same per‑task effective rank, JIVEAdapter matches or outperforms strong single‑task and multi‑task low‑rank baselines. Unlike methods that require extra Mixture‑of‑Experts modules, JIVEAdapter relies solely on the Joint and Individual low‑rank matrices, keeping the architecture simple. When a related held‑in task exists, its frozen Joint can serve a held‑out task by reusing that task’s Individual with only a cheap per‑direction scaling; otherwise a small new Individual is trained.\ \ Review: By explicitly separating shared and task‑specific low‑rank representations, JIVEAdapter achieves both interpretability and parameter efficiency for multi‑task fine‑tuning, and its incremental learning scheme makes it attractive for real‑world deployment.

Original Source: https://arxiv.org/abs/2610.07036

[h] Back to Home