This appendix lists every quantity the paper relies on—its symbol, value, and how we obtained it. Each is measured from the deployed system, derived in closed form from other quantities, fixed by construction, or assumed (with the basis stated). The measured quantities feed the prose and the figures from one shared source, so the two cannot disagree.
Seven measurement procedures produce the data. The build-cost measurement aggregates per-case input/output tokens, USD, and wall-clock from MODEL-stage production logs across 1,184 deployed case versions. The parameter-count audit counts numeric literals in the GWM configuration, counting shared global/sector parameters once and per-case parameters times the universe size. The narrative-arm ablation re-runs the production case drafter as a one-shot NWM over 20 cases under prompt caching, recording cache-adjusted cost, tokens, API calls, and latency. The explanation-quality evaluation scores each NWM memo with a tool-backed judge that recomputes every internal-consistency invariant with a deterministic calculator (yielding and ) and judges hardness-to-vary against Definition 10. The inference-cost benchmark times one belief-propagation pass at samples over recent deployed cases and prices the median warm wall-clock at a standardized on-demand vCPU-hour rate. The predictive-distribution token count measures, over the same scored memos, the tokens of the predictive distribution’s numeric figures (the scenario and direction tables), giving the LLM’s best-case per-PD output. The incremental-PD (re-query) measurement reissues the counterfactual queries of Definition 9 as warm, cache-hit-only calls (cache re-read, no cache write) against the deployed NWM arm over 18 such calls, recording the median cache-read tokens (), median latency, and a median cost (); the pre-blow-up output-token term is then calibrated so that Eq. (8), evaluated at the deployed attainment , reproduces this measured median cost exactly.
Measured.
| Symbol | Quantity | Value | Source / method |
|---|---|---|---|
| GWM build cost per case | Build-cost measurement over deployed case versions; same source fixes build tokens ( in / out) and wall-clock ( min). | ||
| NWM cost per case (sourced memo) | Narrative-arm ablation over cases (range –); same source fixes memo latency ( s). | ||
| GWM cost & latency per predictive distribution | , s | Inference benchmark: median warm belief-propagation wall-clock ( samples) directly gives the latency; a standardized vCPU-hour rate gives the cost. | |
| NWM incremental-PD cached-context and pre-blow-up output tokens | , | Incremental-PD (re-query) measurement over warm do() re-queries: is the median cache-read token count; is calibrated so Eq. (8) reproduces the measured median warm cost () at . Same calls fix the incremental-PD latency ( s). | |
| NWM explanation-quality components | Table 4 | Explanation-quality evaluation over cases (/ tool-checked, judged); best-case PD output tokens via the predictive-distribution token count. | |
| XQ attainment ceilings | Table 4 | bounded joint-correctness probability (Eq. (7)). | |
| subdomain-novelty decay | Heaps’ law (App. C). | ||
| NWM upfront (first memo cached context) | (first memo) cache-write price. | ||
| — | best-case per-PD cost ratio | NWM cached re-read figures-only output (no blow-up) . | |
| — | incremental-PD latency ratio | measured warm NWM re-query latency measured GWM per-PD latency (no best-case extrapolation: latency has a fixed network/decode floor that cost does not, so it is anchored at the directly measured point rather than the cost model’s best-case token count). | |
| — | build break-even | PDs | at XQ target . |
| NWM cost per case (full evidence chain) | full evidence chain. |
Structural (fixed by construction).
| Symbol | Quantity | Value | Basis |
|---|---|---|---|
| GWM components | Table 4 | Exact given evidence: stiffness and counterfactual consistency by construction; is itself the product of rigorous iterative construction (adversarial validation belief propagation empirical calibration), so (bounded misspecification) yet conservatively well above the NWM’s one-shot median (); from exact/large-sample belief propagation. | |
| NWM residual control | Table 4 | Measured per memo: from the simplex volume compatible with the narrative’s probability bands (Appendix C); the GWM’s large-sample is the generous ceiling. |
Assumed (with basis).
| Symbol | Quantity | Value | Basis |
|---|---|---|---|
| token prices ($/Mtok) | Published Opus 4.8 list prices. | ||
| SOTA-LLM parameters | Order of magnitude. | ||
| NWM cold-context tokens; MC-averaged emissions/PD | , | Amortization-model assumptions ( and , the tokens entering Eq. (8) directly, are measured; see the Measured table above). | |
| NWM cost–quality blow-up exponent | Identified, not free: (Eq. (11), Proposition 4); numerical value calibrated (Appendix F). | ||
| subdomain-synthesis tokens; Heaps exponent | Engineering / structural. |