This appendix states the narrative arm’s per-PD cost model used in Section 3, the assumptions behind it, and the resulting blow-up claim; the proof is deferred to Appendix G.
Writing the cached-context re-read as and the per-PD output as scaled by a quality-dependent blow-up factor,
| (8) |
where is the NWM’s structural ceiling (Proposition 5), the tokens to narrate one PD, the cost–quality blow-up exponent, and the number of emissions MC-averaged into one PD.
After Eq. (8), the ceiling is budget-invariant—a structural supremum no spending can exceed (Proposition 5), and for quantification we generously set it equal to the GWM’s, —so it fixes the pole location. What a larger budget does buy is a higher achieved XQ: unstructured validator iteration certifies progressively more of the checkable invariants (the stiffness reconstructions, the static-CC consistency checks, and the HtV-hardening edits), pushing the achieved attainment up toward at a per-PD cost that diverges as . Residual control is improved not by validator iteration but by MC-averaging emissions (the factor above); we grant its large-sample value only as the ceiling, while the one-shot deployed memos attain a measured (Appendix C). The numerical value of is calibrated (Appendix I); the claim below identifies it.
Throughout, the XQ target is set by an ideal adversarial validator that maintains an unbounded battery of checkable invariants—the interventional queries of Definition 9 together with the structural reconstructions of Definitions 8 and 10—and certifies the NWM’s predictive only when it jointly passes a finite subset of size . We derive the two ingredients of the blow-up (coverage and cost) rather than assume them.
Let (, Proposition 5) be the attainment reached when the predictive jointly passes invariants drawn from the battery, and write the residual incorrectness . Because XQ aggregates multiplicatively (Eq. (6)), certifying one further invariant removes a fraction of that is bounded away from , so in the continuum limit
| (9) |
Joint correctness is the product of per-invariant correctness events (Eq. (6)), so the residual is the surviving share of error. An adversary never spends a query on a check already implied by the passed set, and draws invariants of comparable difficulty; hence certifying the -th invariant multiplies the surviving error by a stationary factor with , . Thus , i.e. with ; substituting gives (9). ∎
A narrative WM certifies invariants only by local edits to its state—appended context tokens, retrieved snippets, or gradient steps—none of which is the Bayesian conditioning operator (Proposition 3), and none of which supplies the unified internal mechanism that would satisfy the invariants jointly (Appendix H). Then there is a per-invariant pass rate such that the expected number of refinement passes to jointly certify invariants is bounded below by the crystallization law of Appendix B,
| (10) |
We mirror the measure argument of Proposition 3. (i) Each pass is non-generic. The edit that fixes a freshly drawn invariant maps the state to a new predictive which, by Proposition 3(i), coincides with the Bayes-conditioned target only on a measure-zero subset of states; passing the invariant is therefore a non-generic event of probability at most some . This bound is uniform over the battery: were it not—were some sequence of edits to drive the joint pass rate to across the unbounded battery—then by Richens and Everitt (2024) the NWM would have to embed an approximate causal model and route each intervention to it as a conditioning operation, i.e. its internal state would always implement a GWM even as autoregressive decoding perturbs it. By the non-generic-circuit measure of Proposition 3 (the subset of -parameter circuits implementing exact conditioning has prior measure decreasing in and ) this event does not occur for a mechanism-free , so strictly. (ii) Constraints do not co-crystallize. Because no shared mechanism enforces the invariants jointly, the edit that passes invariant is uncorrelated with those that passed and regresses each with positive probability; the pass-events are thus no better than independent, and the joint-pass (all-certified) set has measure at most . This is exactly the sparse-island structure of Appendix B with binding constraints, whose expected search cost is ; hence . ∎
Combining Lemmas 1 and 2, the NWM per-PD cost follows the output term of Eq. (8) with the exponent identified:
| (11) |
Moreover the pole at is impassable for a categorical, not a budgetary, reason: crossing it requires , which by Richens and Everitt (2024) entails an (approximate) causal mechanism—becoming a GWM—unavailable to a mechanism-free model confined to its function class.
Remark (residual sub-lever). When the predictive is estimated from sampled emissions rather than emitted wholesale, reducing the Monte-Carlo residual adds a second cost scaling as (error , cost linear in )—the special case—which only reinforces the divergence. Proposition 4 thus establishes the output term of Eq. (8) and pins to interpretable quantities; the cost diverges at the structural ceiling, which we generously set equal to the GWM’s (the true ceiling lying strictly below).
Both arms are priced on two consistent bases (Figure 3). On the single-generation basis the GWM pays a one-time build (yielding a reusable posterior) against the NWM’s first run ; on the incremental per-PD basis the GWM pays a flat belief-propagation pass against per query. Cumulative cost after predictive distributions is versus , so the build is recovered after
| (12) |
predictive distributions (Figure 4); falls below one as the quality bar rises and is undefined once exceeds the NWM ceiling, where the NWM cannot reach the target at any budget.