This paper addresses the challenge of inferring 3D cellular biophysics from 2D microscopy when only population‑level statistics—not individual cell labels—are available. We propose a population‑supervised framework that maps each 2D red‑cell image to latent biophysical quantities and aggregates them to obtain mean corpuscular volume (MCV), red‑cell distribution width (RDW), and mean corpuscular haemoglobin (MCH).
The model consists of four key components:
- Shared local inference: a common feature extractor processes all images, capturing local morphological cues;
- Biophysically structured decoder: the decoder outputs are constrained by physical formulas for volume and haemoglobin, ensuring interpretability;
- Learned instance weighting: trainable weights modulate each image’s contribution, handling sample heterogeneity;
- Device‑specific calibration: calibration parameters are introduced for each microscope type to remove cross‑device bias.
We formalise the conditions under which aggregate observations uniquely identify restricted instance predictors and prove that population agreement alone cannot recover single‑cell properties or 3D geometry. Additionally, we derive a dispersion penalty induced by subset‑mean matching, which appears explicitly in the training objective to curb variance inflation of predictions.
Experiments were conducted on a dataset comprising 390 specimens and 1,105 acquisitions across six devices. Compared with a Sysmex analyser, the model achieves Pearson correlations of 0.86–0.98 for MCV, RDW, and MCH, demonstrating that reliable 3D biophysical inference is possible from 2D images and population supervision without explicit 3D reconstruction.
Review: The study presents a novel route to infer single‑cell biophysical traits using only population statistics, cleverly integrating physical constraints with deep learning. It opens new avenues for quantitative microscopy and highlights the potential of population‑level supervision in biomedical imaging.