NeFut Logo NeFut
中 Admin Login

[CS.AI] Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #Neural

In large‑scale learning, training data are usually quantified by a single sample count. For hierarchically and spatially structured data, the same number of samples may come from few or many sources and be distributed differently across the underlying domain. We therefore treat data scaling as an allocation problem, decomposing it into three axes: unique sample count, source diversity, and spatial coverage. Our study uses microscopic whole‑brain histology, where a “source” is an individual brain and a “sample” is an image patch at a specific spatial location. Across 93 controlled pre‑training runs of a contrastive model that leverages spatial proximity for supervision, we vary data allocation, compute budget, and model capacity, processing 11.6 million spatially anchored patches from 21 human brains. Results show performance improves with more unique samples, broader spatial coverage, additional compute, and larger model capacity. When the total sample budget is fixed, distributing samples across 1 to 18 subjects yields no detectable gain, even though representations generalize markedly better to subjects seen during pre‑training. Thus, inter‑subject variation strongly influences generalization, but adding more subjects provides no benefit under a fixed sample budget. The study establishes sample count, source diversity, and spatial coverage as distinct axes of data scaling in spatially structured representation learning.

Review

Original Source: https://arxiv.org/abs/2609.31201

[h] Back to Home