This appendix gives formal definitions for the four components of (Definition 7).
Let be a PD produced by a computation with named structural inputs , a set of invariants (accounting identities, sign and monotonicity constraints), and—for sampling-based WMs—a Monte-Carlo budget yielding standard error on a target summary functional with natural scale . Fix a battery of intervention queries.
Let be the elasticity of to input , and the mechanism-implied admissible band for that elasticity. Then
| (4) |
the fraction of inputs that are not free knobs (responses bounded and mechanism-consistent).
| (5) |
For an alternative outcome , let be the nearest model (minimal structural edit) that fits without violating , and the normalized forced change. Then : a high value means the explanation cannot be cheaply twisted to fit a different outcome (Deutsch, 2011).
, where is the residual uncertainty about the true posterior left by the computation that produced ; in the exact-inference limit. For a sampling WM with Monte-Carlo standard error and tolerance , , which vanishes in the large-sample limit. A single narrative emission instead asserts one scenario distribution while its qualitative rationale pins only a set of distributions on the -scenario probability simplex ; the residual uncertainty is then the linear extent of that set, , so (estimated in Appendix C).
The four components are independent correctness probabilities, so the joint correctness of is their product and explanation quality is the corresponding log-probability:
| (6) |
measured in nats, on the same scale as . Near the ideal, , which recovers to first order the linear KL bound used in the proof of Proposition 1; the log form is its all-orders extension, and unlike an arithmetic mean it admits no cross-component compensation—a single weak axis caps the whole. For figures and quality targets we use the bounded attainment
| (7) |
the joint-correctness probability of the four axes—a strictly monotone transform of (indeed ), so every ceiling and ordering statement transfers between the two, and a ratio of two ’s is exactly the exponential of their difference.
A note on scale, to avoid a common confusion. itself is the log-scale quantity Proposition 1 lower-bounds by, and—exactly as intuition suggests—it increases toward (though it need not reach) as becomes a better explanation, mirroring as ; this is not in tension with attainment being bounded. What lives on the familiar scale is not but . Because the two are a strictly monotone transform of each other, we use them interchangeably in prose below for readability—so wherever the text speaks of “raising XQ,” an “XQ target,” or an “XQ ceiling” expressed as a number in , it denotes this bounded attainment , not the log-scale of Eq. (6), which is what is actually being plotted or bounded in Sections 3.2 and the propositions below.
The definitions above are population quantities over a battery and an invariant set ; we now state how each is estimated on the narrative-arm memos (the measurement of Section 3.1), so the reported scores map transparently onto Definitions 8–11. A tool-backed judge agent (also Claude Opus 4.8) reads each memo equipped with a sandboxed arithmetic evaluator and a numerical closeness check; it may verify a number only by calling a tool, so the structural scores are tool-checked rather than asserted, and the full call log is persisted per case for audit.
Stiffness (Def. 8). The fraction of headline numbers (target price, expected return, direction probabilities, the weighted scenario value, each multiple-to-price bridge) that are reconstructible from other stated inputs in the memo. A number that no stated inputs reproduce is a free knob—an unconstrained elasticity direction, the discrete analogue of .
Counterfactual consistency (Def. 9). Estimated in two steps and combined as , so the weaker step caps the score. Step 1 (static internal consistency): the battery that probabilities sum to one; the stated target equals the probability-weighted scenario prices; expected return target/spot ; the direction buckets match the scenario upsides; scenario prices are monotone in severity; and the recommendation’s sign lies in the stated return band—each an instance of “ satisfies ,” checked with the tools. This step certifies only that the memo is self-consistent, not that it answers correctly. Step 2 (interventional re-query): to test the do-calculus clause of Definition 9 directly, the judge poses at least five off-grid interventions —probes whose correct answer is neither any tabulated scenario price nor a convex re-weighting of them, spanning between-scenario, compound, fixed-input, and mechanism-inversion families—and re-queries the same NWM under each, holding its stated thesis and mechanism fixed. Each re-forecast is scored continuously in for whether it tracks the memo’s own stated mechanism in both direction and magnitude (no external oracle), and is their mean.
Hardness-to-vary (Def. 10). The one component scored as a direct judgment: the minimal structural edit needed to make the rationale fit the opposite recommendation without breaking an invariant, in .
Residual-noise control (Def. 11). Measured per memo. The judge brackets each of the scenario probabilities by a band that the narrative defensibly supports; a deterministic tool computes the fraction of the simplex compatible with those bands (by uniform-simplex sampling), and . A tightly argued distribution gives small and high ; a vague one gives . The GWM’s large-sample is taken as the NWM’s generous ceiling (reachable only by MC-averaging many emissions), not as its deployed value.