Primordia Co.Grounded World Models

Appendix F Cost-Model Details

This appendix states the narrative arm’s per-PD cost model used in Section 3, the assumptions behind it, and the resulting blow-up claim; the proof is deferred to Appendix G.

Setup

Writing the cached-context re-read as Tcachepcache and the per-PD output as nToutpout scaled by a quality-dependent blow-up factor,

cNWM(XQ)=Tcachepcachecached re-read+nToutpout(ANWMANWMXQ)ηoutput (diverges at the ceiling ANWM), (8)

where ANWM is the NWM’s structural ceiling (Proposition 5), Tout the tokens to narrate one PD, η>0 the cost–quality blow-up exponent, and n1 the number of emissions MC-averaged into one PD.

After Eq. (8), the ceiling ANWM is budget-invariant—a structural supremum no spending can exceed (Proposition 5), and for quantification we generously set it equal to the GWM’s, ANWM=AGWM—so it fixes the pole location. What a larger budget does buy is a higher achieved XQ: unstructured validator iteration certifies progressively more of the checkable invariants (the stiffness reconstructions, the static-CC consistency checks, and the HtV-hardening edits), pushing the achieved attainment up toward ANWM at a per-PD cost that diverges as XQANWM. Residual control is improved not by validator iteration but by MC-averaging emissions (the factor n above); we grant its large-sample value only as the ceiling, while the one-shot deployed memos attain a measured R=0.73<0.95 (Appendix C). The numerical value of η is calibrated (Appendix I); the claim below identifies it.

Coverage and the cost of adversarial certification

Throughout, the XQ target is set by an ideal adversarial validator that maintains an unbounded battery of checkable invariants—the interventional queries of Definition 9 together with the structural reconstructions of Definitions 8 and 10—and certifies the NWM’s predictive only when it jointly passes a finite subset of size ν. We derive the two ingredients of the blow-up (coverage and cost) rather than assume them.

Lemma 1 (Saturating coverage).

Let x<xc (xc=ANWM, Proposition 5) be the attainment reached when the predictive jointly passes ν invariants drawn from the battery, and write the residual incorrectness r=1x/xc(0,1). Because XQ aggregates multiplicatively (Eq. (6)), certifying one further invariant removes a fraction of r that is bounded away from 0, so in the continuum limit

dxdν=λ(xcx),λ>0. (9)
Proof.

Joint correctness is the product ici of per-invariant correctness events (Eq. (6)), so the residual r=1x/xc is the surviving share of error. An adversary never spends a query on a check already implied by the passed set, and draws invariants of comparable difficulty; hence certifying the (ν+1)-th invariant multiplies the surviving error by a stationary factor 1β with β(β0,1), β0>0. Thus r(ν+1)=(1β)r(ν), i.e. dr/dν=λr with λ=ln(1β)>0; substituting r=1x/xc gives (9). ∎

Lemma 2 (Adversarial certification is geometrically hard for a mechanism-free model).

A narrative WM certifies invariants only by local edits to its state—appended context tokens, retrieved snippets, or gradient steps—none of which is the Bayesian conditioning operator Be (Proposition 3), and none of which supplies the unified internal mechanism that would satisfy the invariants jointly (Appendix H). Then there is a per-invariant pass rate q¯<1 such that the expected number of O(1) refinement passes to jointly certify ν invariants is bounded below by the crystallization law of Appendix B,

k(ν)q¯ν. (10)
Proof.

We mirror the measure argument of Proposition 3. (i) Each pass is non-generic. The edit that fixes a freshly drawn invariant maps the state to a new predictive which, by Proposition 3(i), coincides with the Bayes-conditioned target only on a measure-zero subset of states; passing the invariant is therefore a non-generic event of probability at most some q¯<1. This bound is uniform over the battery: were it not—were some sequence of edits to drive the joint pass rate to 1 across the unbounded battery—then by Richens and Everitt (2024) the NWM would have to embed an approximate causal model and route each intervention to it as a conditioning operation, i.e. its internal state would always implement a GWM even as autoregressive decoding perturbs it. By the non-generic-circuit measure of Proposition 3 (the subset of |L|-parameter circuits implementing exact conditioning has prior measure decreasing in |L| and K) this event does not occur for a mechanism-free L, so q¯<1 strictly. (ii) Constraints do not co-crystallize. Because no shared mechanism enforces the invariants jointly, the edit that passes invariant ν+1 is uncorrelated with those that passed 1,,ν and regresses each with positive probability; the ν pass-events are thus no better than independent, and the joint-pass (all-certified) set has measure at most q¯ν. This is exactly the sparse-island structure of Appendix B with m=ν binding constraints, whose expected search cost is q¯m; hence k(ν)q¯ν. ∎

Claim

Proposition 4 (NWM cost blow-up).

Combining Lemmas 1 and 2, the NWM per-PD cost follows the output term of Eq. (8) with the exponent identified:

C(x)q¯ν(x)=(xcxcx)η,η=ln(1/q¯)λ> 0. (11)

Moreover the pole at xc is impassable for a categorical, not a budgetary, reason: crossing it requires CC1, which by Richens and Everitt (2024) entails an (approximate) causal mechanism—becoming a GWM—unavailable to a mechanism-free model confined to its function class.

Remark (residual sub-lever). When the predictive is estimated from n sampled emissions rather than emitted wholesale, reducing the Monte-Carlo residual adds a second cost scaling as (1R)2 (error εn1/2, cost linear in n)—the η=2 special case—which only reinforces the divergence. Proposition 4 thus establishes the output term of Eq. (8) and pins η to interpretable quantities; the cost diverges at the structural ceiling, which we generously set equal to the GWM’s (the true ceiling lying strictly below).

Pricing bases and break-even

Both arms are priced on two consistent bases (Figure 3). On the single-generation basis the GWM pays a one-time build cbuild (yielding a reusable posterior) against the NWM’s first run cNWM+cNWM(XQ); on the incremental per-PD basis the GWM pays a flat belief-propagation pass cPD against cNWM(XQ) per query. Cumulative cost after Q predictive distributions is CGWM(Q)=cbuild+QcPD versus CNWM(Q)=cNWM+QcNWM(XQ), so the build is recovered after

Q(XQ)=cbuildcNWMcNWM(XQ)cPD (12)

predictive distributions (Figure 4); Q falls below one as the quality bar rises and is undefined once XQ exceeds the NWM ceiling, where the NWM cannot reach the target at any budget.