Vision‑language‑action (VLA) policies link a pretrained vision‑language backbone to an action head via a latent interface, yet it is unclear which backbone layers should be exposed. We evaluate single‑layer selection and multi‑layer fusion on frozen backbones across three pretrained models and two manipulation benchmarks (LIBERO, CALVIN), using three policy‑training seeds per configuration. Across three fusion mechanisms and three layer‑subset strategies, 47 of 54 configurations underperform the best observed single‑layer policy. Statistical analysis confirms that the advantage of fusion is very limited. The optimal layer varies substantially across backbones and benchmarks, making exhaustive policy sweeps costly. We derive a reweighting equivalence between the information‑bottleneck objectives for action‑conditioned InfoNCE and action prediction, motivating InfoNCE as a proxy for layer quality. Empirically, InfoNCE shows the most consistent positive association with policy success among four evaluated proxies. Selecting the layer with the highest InfoNCE score reduces GPU compute by 9‑33× compared to exhaustive sweeps and lowers mean selection regret from 17.89 percentage points (deepest‑layer selection) to 3.71 points across six settings. Its mean regret approaches the 3.28‑3.50 points achieved by fixed‑layer heuristics optimized retrospectively using all six oracle sweeps, without requiring closed‑loop evaluations during selection.
Review: This study highlights the importance of layer choice in VLA systems and offers an InfoNCE‑based selection strategy that dramatically cuts computational cost while maintaining policy performance, providing a practical solution for real‑world deployment.