Primordia Co.Grounded World Models

Appendix E A Worked Example: GWM vs NWM on One Case

This appendix grounds the protocol of Section 3.1 in one real case from the sample—anonymized as ACME, spot $254.34—showing both the single-PD quality contrast and the cost of a stream of PDs. All numbers are taken from the released artifacts: the NWM memo and its tool-checked judge log, and the GWM’s evaluated posterior.

One PD: the quality contrast

Both arms received the same dossier and fundamentals and emit the same schema. Table 2 places the two outputs side by side.

NWM (asserted) GWM (computed)
Scenario Prob Price Upside Prob Price Return
WORST 5% $150 41.0% 2.6% $0 100%
BAD 20% $215 15.5% 18.8% $184.0 27.7%
BASE 40% $298 +17.2% 15.1% $255.0 +0.2%
GOOD 25% $360 +41.5% 46.1% $368.7 +45.0%
BEST 10% $430 +69.1% 17.4% $608.5 +139.3%
Target price $298 $289.7
Expected return +17.2% +13.9%
Direction L/H/S 65 / 30 / 5 (table implies 75 / 20 / 5) 63.5 / 26.8 / 9.7
Recommendation LONG LONG
Conviction Medium HIGH
Table 2: ACME, side by side: asserted vs computed (spot $254.34). Red = asserted by the model; blue = computed from the posterior (read off the same Monte-Carlo samples). The NWM’s direction probabilities (65/30/5) are a free knob that contradicts its own scenario table, which implies 75/20/5; the GWM’s (63.5/26.8/9.7) are Pr(upside>hurdle) over the scenario samples and so cannot disagree with the table. The GWM target is the probability-weighted geometric-mean certainty-equivalent (vs the NWM’s arithmetic weighted price), and its WORST bucket is a sub-78.8% wipeout tail.

What the judge finds.

The tool-backed judge (Appendix C) reconstructs the NWM’s headline numbers with its calculator: the target reconstructs as the probability-weighted price (ipiPi=302.7298, within 2%), expected return as 298/254.341=+17.2%, and each scenario upside as Pi/254.341—all pass. One number does not reconstruct: the asserted direction probabilities 65/30/5 contradict the scenario table, which implies 40+25+10=75% LONG, 20% HOLD, 5% SHORT. This single free knob fails one of nine static consistency checks (CCstat=8/9=0.889) and one of seven stiffness checks (S=6/7=0.857). Probing further, the judge re-queried the same NWM under five off-grid interventions—a between-scenario multiple compression, a buyback that accretes EPS, a compound revenue/margin shift, an interpolated revenue/margin/multiple triple, and a breakeven-multiple inversion, none answerable from the scenario table—and every re-forecast tracked the memo’s own EPS×multiple bridge in direction and magnitude (CCint=1.0), so CC=min(0.889,1.0)=0.889. Hardness-to-vary is judged 0.62 (the soft scenario probabilities and a reversible narrative can be re-spun for a bear case without breaking the arithmetic). Residual control is measured from the probability bands the narrative supports: only f=0.8% of the five-scenario simplex is compatible, giving R=1f1/4=0.705—high, but below the GWM’s large-sample 0.95.

The GWM has no such knob to break. Its scenario probabilities sum to one, each scenario upside equals Pi/spot1, prices are monotone in severity, and—crucially—its direction probabilities are not asserted but read off the same posterior samples as the scenarios, so they cannot disagree with the table. Stiffness and counterfactual consistency are therefore 1 by construction (Table 4); the residual gap to a perfect explanation is bounded misspecification (HtV=0.92) and Monte-Carlo residual (R=0.95), not a free parameter. Table 3 collects the scores.

Component NWM (ACME) NWM (median) GWM
Stiffness S 0.857 0.83 1
Counterfactual consistency CC 0.889 0.88 1
Hardness-to-vary HtV 0.62 0.67 0.92
Residual control R 0.705 0.73 0.95
Attainment A=ici 0.36 0.87
Table 3: Explanation-quality scorecard for the worked case. ACME’s measured (S,CC,HtV,R) are near the sample medians, yet still bounded well away from the GWM’s S=CC=1. Because XQ aggregates multiplicatively (Eq. (6)), the NWM’s joint correctness trails the GWM by 2.4×.

Across the stream.

Table 3 is the Q=1 slice, which most flatters the NWM. ACME is in fact queried repeatedly—each what-if and re-forecast is another PD—so its per-case economics are those of Sections 3.23.2: ACME’s GWM build is paid once, and every follow-on PD, including the do(a) battery (re-rating the exit multiple, shifting demand, flipping ValuationReRates), is a flat cPD call that preserves the invariants above (CC1). A one-shot NWM run costs up to $1.82 and buys no reusable posterior (only a KV cache of the run’s input and output tokens), so every additional PD is another run at a cost that climbs with the XQ bar. The crossover and ceilings are exactly those of Figure 4.