Primordia Co.Grounded World Models

Appendix G Proofs

Throughout, qE denotes a GWM’s posterior predictive for Y after conditioning on evidence E (Eq. (1)), and p the reference of Definition 6. Grounding is the following conjugate Beta–Bernoulli update:

p(hE)=Beta(αh,βh),αh=1+eEh+we,βh=1+eEhwe, (13)

where we is the trust weight of source e and Eh+,Eh are supporting and refuting evidence Hence, at the level of each evidentiary hypothesis the model is well-specified and the update is exact Bayesian conditioning.

The results below use two assumptions, stated here (the counterfactual battery 𝒜 and invariant set are formalized in Appendix C).

Assumption 1 (Representative checks).

The finite battery 𝒜 of counterfactual queries and the invariant set are representative of the query/outcome distribution of interest, so that certifying them bounds the error on untested queries up to an O(ϵ) remainder.

Assumption 2 (Bounded misspecification).

The GWM is not grossly misspecified: p (or its best computable approximation, in the M-open case) lies within the support of the model class entertained by g.212121Checking whether this assumption holds—comparing the model against held-out outcomes, revising its structure when it fails—is itself not a coherent Bayesian operation: it requires stepping outside the model class entertained by g to entertain structures the prior never covered, the “Cantor’s corner” tension identified by Gelman and Yao (2021). We do not claim Assumption 2 is verified from within the formalism; we claim it is checked and iteratively repaired by the program-synthesis and validation loop described in Section 1, a practice external to, and in that sense in tension with, strict Bayesian coherence.

Empirical support. These realizability conditions are supported by confidential quantitative and qualitative data from forward-testing and daily system usage (internally since December 2025, publicly live since February 2026). A follow-up paper will provide public evidence that the class entertained by g contains a sufficiently accurate, recoverable causal model of the target domain, i.e. that realizability binds in practice and not merely in principle.

Lemma 3 (Prequential information bound).

Let evidence Et={e1,,et} accrue under p, and suppose the prior assigns mass π>0 to the hypothesis realized by p (Assumption 2). Then the Bayes predictive losses satisfy

t1𝔼pDKL(p(Et1)qEt1)logπ<. (14)

In particular the expected per-step divergence is summable, hence 0, and cannot persistently increase.

Proof.

The cumulative log-loss of the Bayes mixture predictor telescopes to log of the marginal likelihood of the data, and the marginal likelihood is bounded below by π times the truth’s likelihood (retain only the true component of the mixture). Taking expectations under p and applying the chain rule of relative entropy yields (14); this is the standard prequential/MDL redundancy bound for Bayesian mixtures (Hoeting et al., 1999; Solomonoff, 1964). Summability of the non-negative terms forces them to 0. ∎

Proof of Proposition 2 (Bayesian quality guarantee)

Proof.

Finiteness. By Assumption 2, p (or its best computable approximation, in the M-open case) lies in the support of the class entertained by g, so qE is absolutely continuous with respect to p on the relevant support and DKL(pqE)<.

Monotonicity. For the conjugate update (13), the posterior predictive is the Bayes-optimal predictor under log-loss. By Lemma 3 the expected divergence is summable and cannot persistently increase; for a single well-specified conjugate family it is monotone non-increasing at each step, since absorbing e only sharpens Beta(αh,βh) toward the realized frequency. Hence 𝔼DKL(pqEe)DKL(pqE).

No narrative analogue. An NWM emits q~ directly (or estimates it from samples) with no conditioning operator: “updating” is re-prompting the language model L, a map not constrained to be Bayesian, so DKL(pq~) may increase with new evidence; and when the predictive is read from sampled points, its law can depend on prompt framing, so the divergence need not even be well-defined. It does not inherit the guarantee. ∎

Proof of Proposition 3 (narrative updating is improper)

Proof.

Model the NWM as a fixed language model L equipped with an update operator UL that maps a state—context tokens, or a knowledge base together with a retrieval policy—and a new evidence item e to a new state, from which the implied predictive q~ is read by decoding. Write Be for the exact Bayesian update p(E)p(Ee).

(i) Non-preservation. UL is computed by attention over a finite, compressed context (or by a bounded-recall retrieval step), and its output is the decoded next-token law—a continuous function of the prompt embedding with no constraint tying it to Be. The two maps therefore agree only on a measure-zero subset of (L,e): generically q~Ee=UL(q~E,e)Be(p(E))=p(Ee) even when q~E=p(E). The discrepancy is governed by attention allocation, context-window truncation and compression, and retrieval precision, none of which implement conditioning, so propriety is not invariant under UL.

(ii) Non-convergence and unbounded error. Because the per-step map is not a Bayes update, the cumulative log-loss does not telescope to log of the marginal likelihood, the prequential bound (14) does not apply, and nothing drives the per-step divergence to zero. Concretely, the context window is finite: for any horizon there is an evidence stream whose informative items are evicted or compressed away, after which q~ is independent of them. Choosing a query Y whose answer depends on the evicted evidence makes p(E) place mass where q~ places none, so DKL(pq~) exceeds any prescribed δ. The same holds for KB-augmentation whenever the retrieval policy misses the relevant item.

The exception. Both failures vanish only if L internally represents the exact mechanism of p and its attention/retrieval routes each e to that representation as a conditioning operation. Treating the realized circuit as a draw from the function class of an |L|-parameter network of structural complexity K, the subset implementing exact conditioning has prior measure decreasing in both K and |L|: more parameters admit exponentially more circuits, among which the correctly-wired one is a vanishing fraction. The event is thus non-generic and becomes less likely as either grows.

Consequences. Two consequences are stated formally in Appendix H: enlarging L does not raise the prior mass of the exception (Corollary 2), and replacing UL with a gradient step does not recover propriety either (Corollary 3). In both cases a guarantee of a proper posterior requires an explicit, verified conditioning operator external to L—coupling L to a GWM it calls (Section 4.1). ∎

Proposition 1 (XQ proxies PQ), full statement

Write the excess risk DKL(pq) as approximation plus estimation error. Then, on the queries in the battery 𝒜 and with no further assumption: (i) the residual component R bounds the estimation term (it is the Monte-Carlo error, zero in the exact-inference limit); (ii) the counterfactual-consistency component CC bounds the approximation error on the tested interventions (satisfying the do-calculus identities is exactness on those queries); and (iii) under a Lipschitz regularity condition the stiffness component S bounds the estimation term via feasible-set complexity. Consequently XQ is monotonically related to PQ on 𝒜. Under Assumption 1 this extends to all queries (with HtV controlling the extrapolation) and under Assumption 2 the divergence is finite; increasing XQ then cannot decrease PQ beyond an O(ϵ) slack.

We first isolate the contributions that hold unconditionally, then add the single assumption needed for extrapolation; this makes precise the sense in which mechanistic faithfulness is mostly a result.

Decompose the excess risk into approximation and estimation terms,

DKL(pq)=DKL(pq)approximation+𝔼p[logq/q]estimation, (15)

where q is the best predictive attainable within the model’s structural constraints.

Lemma 4 (Unconditional component bounds).

Without any faithfulness assumption:

  1. (a)

    (Residual.) The estimation term equals the Monte-Carlo discrepancy between the sampled q and the exact predictive q; by the delta method it is O(se(q)2)=O((1R)2), vanishing as R1.

  2. (b)

    (Counterfactual.) On each tested intervention a𝒜, CC=1 means q(do(a)) satisfies the do-calculus identities that p also satisfies; hence the approximation term restricted to 𝒜 is zero, and in general is O(1CC).

  3. (c)

    (Stiffness.) If the target functional T is L-Lipschitz in logθ, the estimation error is bounded by L2 times the volume of the elasticity-feasible set, which is O(1S) (each free knob adds one unconstrained direction).

Proof.

(a) is the standard delta-method variance of a smooth functional of a Monte-Carlo estimate. (b) is immediate from Pearl’s do-calculus: matching the identities is equality of the interventional distributions on 𝒜, so the KL contribution there is 0; the linear-in-(1CC) bound follows by counting violated checks. (c) is a covering-number bound: with dfree=m(1S) unconstrained directions and an L-Lipschitz T, the estimation variance scales with the feasible-set volume dfree. ∎

Proof of Proposition 1.

Lemma 4 already gives, on the checked queries, DKL(pq)c1(1CC)+c2(1S)+c3(1R)2 for constants ci depending on L and the battery—no faithfulness assumption used. Thus higher CC, S, R provably lower the divergence on 𝒜, so on those queries XQ is monotonically related to PQ.

It remains to pass from the checked queries to the full query distribution. By Assumption 1 (representative checks), the un-tested approximation error is at most an O(ϵ) remainder, and HtV controls it: a high HtV(q) means few alternative models fit the same data, so by the Occam/MDL argument the certified-on-𝒜 model is close to p off 𝒜 as well (the surviving-explanation volume is small). Assumption 2 keeps the divergence finite. Combining, DKL(pq)c1(1CC)+c2(1S)+c3(1R)2+c4(1HtV)+O(ϵ). This bound is a sum of gaps (1), whereas XQ of (6) is a sum of logs; the elementary inequality 1xlogx on (0,1] (and (1R)21RlogR) bounds each gap by the negative log of its attainability, so with c¯=max{c1,c2,c3,c4},

PQ(q)=DKL(pq)c¯XQ(q)O(ϵ), (16)

a monotone affine lower bound, tight to first order as the attainabilities approach 1. This one-sided relation—not a bijection—is exactly what the downstream results require: XQ lower-bounds closeness to p, so raising XQ raises the guaranteed floor on PQ, and XQ can never certify a prediction quality the model does not have. Equal-XQ models may still differ in PQ, which the proposition’s “up to a bounded slack” already permits. Hence XQ is a valid observable proxy for the unmeasurable PQ, and the CC, R, S channels are results rather than assumptions. ∎

XQ ceilings, deployed attainment, and the quality gap

Proof.

We must distinguish two quantities. The ceiling of a WM class is the attainable supremum of its components, aggregated by (6); the deployed attainment is the joint correctness the model actually realizes at its operating budget—a point on the cost curve. For a GWM the two coincide: a posterior is exact given its evidence, so the GWM sits at its ceiling at flat cost (attained = ceiling). Being mechanism-pinned it is exact-by-construction on stiffness (S=1) and counterfactual consistency (CC=1); its hardness-to-vary is <1, reflecting bounded misspecification (Assumption 2); and it has the highest residual-noise control R, because exact / large-sample belief propagation drives se(q) to near zero. For a NWM the two differ: the table reports its deployed attainment, with the structural triple (S,CC,HtV) taken as the measured medians over the one-shot NWM memos (Section 3.2). Residual control is also measured at the deployed budget (the simplex-volume residual-uncertainty of Definition 11): the NWM’s one-shot R=0.73 sits below the GWM’s large-sample R=0.95, so R does not cancel—the GWM leads on all four axes. The GWM’s R is granted only as the NWM’s theoretical ceiling (reachable by MC-averaging many emissions). The measured CC is the two-step estimator of Appendix C, min(CCstat,CCint), combining the static internal-consistency battery with an off-grid interventional re-query of the same NWM; at the deployed budget the static step binds, so the reported CC remains a conservative estimate of the interventional ideal of Definition 9. The NWM’s theoretical ceiling is strictly below the GWM’s (Proposition 5); for all quantification we conservatively set it equal to the GWM’s, so the cost-divergence and dominance results hold a fortiori. The per-component values are:

  • GWM (attained = ceiling) (S,CC,HtV,R)=(1,1,0.92,0.95). S=CC=1 because the probabilistic program enforces the elasticity bands and the do-calculus identities exactly (Definitions 8 and 9); HtV=0.92<1 encodes bounded misspecification (Assumption 2)—conditioning still leaves some explanatory slack; and R=0.95 is the highest residual control, from exact / large-sample belief propagation (Definition 11).

  • NWM (deployed, one-shot) (0.83,0.88,0.67,0.73). All four are measured medians over the one-shot NWM memos—a deployed attainment, not a ceiling; the fourth, R=0.73, is the simplex-volume residual control of Definition 11 and sits below the GWM’s R=0.95, so the two arms differ on all four axes (the GWM’s R is granted only as the NWM’s theoretical ceiling).

Aggregating by Eq. (6) gives, for each arm, the log-XQ (nats) and the bounded attainment A=ici=exp(XQ) used as the target axis in the figures—which for the GWM is its ceiling and for the NWM is its deployed attainment. The deployed-budget quality gap is the ratio of the two attainments, which equals 2.4.

Class S CC HtV R XQ (nats) A=ici
GWM 1 1 0.92 0.95 0.13 0.87
NWM (deployed) 0.83 0.88 0.67 0.73 1.02 0.36
Table 4: Per-component values and the aggregate XQ of Eq. (6): the log-XQ ilogci (nats) and the bounded attainability A=ici=exp(XQ)—the joint-correctness probability—used as the figure axis (all held in / derived from the shared parameter set; NWM structural components measured over 20 cases by the explanation-quality evaluation). The NWM row reports its attainment at the deployed one-shot budget; its theoretical ceiling is generously taken equal to the GWM’s (A=0.87). Multiplicative aggregation forbids cross-axis compensation: at the deployed budget the GWM’s joint correctness (0.87) exceeds the NWM’s (0.36) by 2.4×.

Two ceilings for narrative WMs, and GWM dominance

A narrative WM faces two distinct XQ levels: a structural ceiling that no budget can clear (which we generously equate to the GWM’s, Proposition 5), and an earlier-binding level it actually attains under finite compute, set by the cost model of Appendix F. The structural-ceiling existence (Proposition 5) is unconditional; the measured deployed-budget dominance (Corollary 1) is empirical; and the budget-binding level (Proposition 6) depends on the assumed cost model.

Proposition 5 (Structural (theoretical) XQ ceiling).

A narrative WM, lacking an explicit causal mechanism, has structural components bounded away from 1: there exist S¯,CC¯,HtV¯<1 with SS¯, CCCC¯, HtVHtV¯ at every budget. Granting the NWM the GWM’s large-sample residual control R¯=RGWM as its ceiling (generously: MC-averaging many emissions can drive R to that value, though the deployed one-shot R is lower),

XQNWMXQ¯=logS¯+logCC¯+logHtV¯+logR¯< 0, (17)

equivalently A¯NWM<1 for the bounded attainment of Eq. (7). The bound is structural: set entirely by the mechanism-free triple (S,CC,HtV) and holding at every budget.

Proof.

Each structural component certifies a property a mechanism-free model cannot guarantee. (CC) Without a do-operator the interventional distribution is not computed from a causal graph, so there exist interventions a for which the do-calculus identities fail and the pass fraction is CC¯<1. This is not merely an artifact of the present construction: by Richens and Everitt (2024), robustly passing a large interventional battery requires an approximate causal model, so a model that achieved CC1 would have implicitly learned one—and would thereby be a GWM, contradicting the premise that the narrative WM carries no explicit mechanism. (S) Without named structural parameters carrying admissible elasticity bands, at least one input acts as a free knob, so S=1dfree/mS¯<1. (HtV) A free-form rationale admits a local edit fitting an alternative outcome without breaking any checkable invariant, so δ(y)<1 on a positive-measure set and HtVHtV¯<1. Each bound is independent of compute. Substituting the suprema into (6) gives XQ¯<0 (a sum of logs of quantities <1). Residual control is reducible to the GWM’s large-sample value by MC-averaging—whether the predictive is drawn as samples or emitted wholesale—so it does not affect the structural bound. ∎

Remark (what we measure, and a generous convention). The numbers in Table 4 are not the suprema S¯,CC¯,HtV¯ of this proposition; they are the medians attained by the deployed one-shot NWM, i.e. a point at its deployed budget, generally below the suprema. Pinning the suprema numerically is unnecessary for our conclusions: this proposition guarantees A¯NWM<AGWM, but for all cost and dominance quantification we generously set the NWM’s ceiling xc equal to the GWM’s, xc=AGWM. Since the true ceiling is strictly lower, every divergence and dominance statement holds a fortiori; the empirical gap we report (Corollary 1) is then the conservative deployed-budget gap, not an inflated ceiling gap.

Proposition 6 (Scaling (budget-binding) XQ ceiling).

Let xA denote the bounded attainment (Eq. (7)), the [0,1] image of XQ used as the cost-target axis, and assume the cost model of Eq. (8) (Appendix F), whose output term near the ceiling reads C(x)=c0(xc/(xcx))η with η>0 and xc=A¯ the structural attainment ceiling of Proposition 5. Then under any finite budget B the attained value is

x(B)=xc(1(c0/B)1/η)<xc, (18)

strictly below xc and rising to it only as B. The practical ceiling x(B) therefore binds earlier (at lower XQ) than the structural ceiling.

Proof.

Inverting C(x)=B gives (xc/(xcx))η=B/c0, so xcx=xc(c0/B)1/η and x(B)=xc(1(c0/B)1/η). Since c0,B,η>0 the correction is positive, so x(B)<xc, 0 as B. The divergence of C as xxc is established by the cost-model blow-up result (Proposition 4, stated in Appendix F and proved below), which derives the output term of Eq. (8) and identifies η. ∎

Corollary 1 (GWM dominance at the deployed budget).

Because XQ aggregates multiplicatively (Eq. (6)), there is no cross-axis compensation. At the NWM’s deployed (one-shot) budget the measured attainments satisfy SG>SN, CCG>CCN, HtVG>HtVN, and RG>RN (all four measured), so

XQGWMXQNWM=ilogci,Gci,N> 0. (19)

The GWM attains its value at flat cost, whereas raising the NWM’s achieved XQ toward the (generously shared) ceiling costs divergently (Propositions 64); the deployed-budget gap therefore closes only as NWM cost , so the GWM dominates at every finite budget.

Proof.

Every factor ci,G/ci,N1 and the three structural ones are >1 at the deployed budget, so the sum in (19) is positive; equivalently the joint-correctness ratio ici,G/ici,N>1. With the values of Table 4 the GWM’s joint correctness ici,G=0.87 exceeds the NWM’s 0.36 by 2.4×. Closing this gap requires raising the NWM’s achieved XQ, whose per-PD cost diverges as xxc (Proposition 6), while the GWM holds its value at the flat cost cPD; hence at any finite budget the cost-equalized comparison strictly favors the GWM. ∎

Remark. Multiplicative aggregation is what makes the dominance robust: under the earlier arithmetic mean a single strong axis could mask a weak one, but a product is capped by its weakest factor. Since the GWM strictly dominates on all four axes at the deployed budget—including residual control—no mechanism-free narrative can match it without divergent spend.

Cost-model blow-up

Proof of Proposition 4.

Integrating the saturating-coverage law of Lemma 1, dx/(xcx)=λdν, yields ln(xcx)=λν+const, hence

ν(x)=1λlnxcxcx,

which diverges as xxc. By Lemma 2 the refinement cost is C(x)kq¯ν(x)=exp(ν(x)ln(1/q¯))=(xc/(xcx))η with η=ln(1/q¯)/λ>0, which is Eq. (11). Since Lemma 2 bounds the pass count from below by q¯ν, this is a lower bound on cost (equivalently, an upper bound on the attainment a fixed budget buys), so the leading-order blow-up is as stated. Finally the pole is impassable: by Proposition 5 the structural components are bounded away from 1 at every budget, and reaching CC1 would by Richens and Everitt (2024) require an approximate causal model, contradicting the mechanism-free premise; hence x<xc at any finite budget and C as xxc. ∎