This work conducts a systematic comparison of four pre‑training regimes: no pre‑training, class‑disjoint in‑domain pre‑training, supervised out‑of‑domain pre‑training, and label‑free out‑of‑domain pre‑training. Experiments span eight datasets, three few‑shot architectures, and various way‑shot configurations. The findings reveal that class disjointness alone does not eliminate target‑domain influence. In‑domain pre‑training improves over the no‑pre‑training baseline by an average of 33.41 percentage points, while supervised out‑of‑domain pre‑training yields a 23.75‑point gain, indicating a 9.66‑point optimistic bias due to domain overlap. Although out‑of‑domain pre‑training better reflects scenarios with scarce target data, its success strongly depends on source‑target domain compatibility. Moreover, labeled source data are not strictly required: an augmentation‑based label‑free strategy gains 27.71 percentage points on average, closely matching supervised out‑of‑domain pre‑training’s 27.97 points. The authors also introduce a descriptor‑based source‑selection technique that estimates source suitability before pre‑training, achieving a median gap of only 1.37 points from an oracle selector. In sum, the study urges moving beyond in‑domain pre‑training as the default few‑shot evaluation protocol, as it can overestimate performance in realistic low‑data settings.
Review