Scaling studies for AI increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. A higher score under a larger budget alone does not reveal where extra resources should be spent. This review compares evidence across pretraining, test‑time computation, retrieval, and agent evaluation, separating the performance of a tested procedure from the best performance theoretically achievable under the same resource limit. The synthesis identifies three recurring mismatches: counting success before an answer is chosen, using information unavailable to a deployed system, and omitting costs from the comparison. To address these, a capability surface is introduced, expressing performance as a function of budget, mechanism, and available information. Analytical examples show how the evaluation metric, deployment volume, selection rule, and stopping policy can change the allocation conclusion. A resource envelope framework is then proposed to record the task, development and run‑time resources, information access, and the exact procedure behind a reported score. Applying the envelope to a published comparison clarifies which conclusions are supported by evidence and which deployment questions remain open. The resulting framework specifies the necessary comparisons for choosing among feasible systems and motivates experiments on transferring allocation rules across tasks and operating conditions. It does not claim a universal scaling law nor infer general intelligence from benchmark gains.
Review