Capability deep-dive · portfolio stress
The board asks what the tail costs. You hand over one number with no way to reproduce it.
The spreadsheet that produced it has been edited since, the assumptions live in a different file, and the person who ran it has moved teams. Nobody doubts the number because nobody can check it — which is a different thing from the number being right.
Reproducibility is the claim. The result is not
The engine facts are read from our own module: a deterministic generator, Box–Muller draws, a correlation matrix with a Cholesky factor, nearest-rank percentiles, and a hard floor on iterations with a verbose refusal below it. The back-test power figure is our own derivation, not an external benchmark: at a one-in-two-hundred annual exceedance over ten years, a test that rejects on a single exceedance has a size just under five per cent and just under ten per cent power against a model that is wrong by a factor of two. We print it because a stress architecture that lets its accuracy table imply it validates the capital point is asserting the absence of evidence as evidence.
SAME SEED, SAME RUN, OR THE DOSSIER REFUSES
A run whose seed does not reproduce is inadmissible, and the dossier compiler refuses to cite it. That is the inversion this page rests on: the headline claim is about reproducibility, not about results. A distribution looks like evidence, which makes a chart the most dangerous object on a marketing site — so where a shape is drawn here at all it is drawn as a shape, every displayed value marked as uncalibrated inside its own frame.
An immutable snapshot plus a lattice of reversible deltas
Reversibility is in the object graph, not in a flag
A scenario is a row that references an immutable base snapshot; its results land in their own table; deletion is soft and preserves results. The base is never mutated, so a stress run cannot damage the thing it was run against. Nobody has to remember to set anything.
A scenario on an unfinished snapshot is refused
A snapshot that is still being built has an empty state, and a scenario over it would produce, in the source’s own words, garbage results that look like a legitimate outcome. That case throws today rather than computing. It is the shipped refusal this whole architecture is an extension of.
The engine has no domain in it
The engine’s entire coupling to its subject is one function type: samples in, a number out. Everything else is domain-free. Swapping a revenue driver for a loss driver is a swap of that one function, not a rewrite — which is why the maths is reusable and why nothing about the maths makes a catastrophe model appear.
A scalar delta is not a capital answer
The quantity a stress cycle reports is a distribution, and the delta between two distributions is not the difference of their means. The honest minimum is a percentile band with its tail support — a band rendered without the number of tail samples behind it is false precision, and the comparison surface is built so that it cannot be.
seeded correlated Monte-Carlo(shipped)ROS/src/services/gtm-genesis/blueprint/MonteCarloEngine.tssnapshots, scenarios, comparisons, back-test store(shipped)ROS/src/services/DigitalTwinService.ts
What a back-test can and cannot validate
Insurance accuracy is not a time series, because the label arrives late and arrives continuously. It is a matrix of origin cohort against maturity lag, and a per-period accuracy row is a diagonal of that matrix — which is precisely the diagonal that mixes immature and mature cohorts into one number. What belongs in the cells is a distributional test on the probability integral transform — a standard check that predicted distributions match the outcomes that arrived — and a score in loss units for anything a committee will read, rather than a percentage error on a capital forecast.
And the bound: nothing in a back-test validates a one-in-two-hundred. The table validates the body of the distribution, the frequency, the vulnerability by cohort and the aggregate — and it has to say so on its own face. A calibration entry also has to become an entry rather than an edit, because a back-test needs to know which parameter value was live for a cohort at the time, and today calibration overwrites.
Where a stress run fails closed
REFUSED
A replay of this run under its recorded seed produced a different result. The run is inadmissible and the dossier compiler refused to cite it.
A capital figure that cannot be reproduced is an assertion, and an assertion is exactly what a validation regime is trying to eliminate. The seed as submitted is not the run’s identity — the engine coerces it before use, so two different submissions can collapse into one random stream — so the record stores the coerced seed with the digests (fingerprints) of the exact code and toolchain beside it. Two neighbouring refusals share the shape: a run with no deadline result is a loud machine-coded refusal rather than an indefinite wait or a silently substituted prior run, and a comparison across a changed driver set refuses rather than differencing two incomparable objects.
Details for support
CHAIN_DIVERGENCE
What this connects to
Validation evidence
Where a reproducible run becomes a dossier — and the refusal that stops an irreproducible one being cited.
Deliberation rooms
Who signs a capital statement, what a substantive sign-off records, and the measurement that would catch a ceremonial one.
Four working papers. A person sends the one you pick — no download wall, and no meeting is booked.
Work out what this costs