Outcome · Deployment readiness
The model works. It has sat in staging for months because nobody is able to sign for it.
The blocker is rarely the science. It is that promotion needs an owner willing to put a name to it, an evaluation somebody else can re-run, and a record of what was compared against what — all assembled afterwards, by the team with the least appetite for assembling it.
The gates come first, and there are exactly three answers to failing one
How we count this
All three figures are the demo-data corpus specification’s own — specified and gate-carried, never “measured on our corpus”, because the corpus is a specification, not generated data. The response that is missing from the list of three is the one every green build is usually bought with: weakening a gate to make a build pass. It is not an option here, by construction, and its absence is the point of printing the list.The run the compiler refuses, on purpose
One run in the specification is deliberately irreproducible — an environment pin that was never captured. The dossier compiler refuses it and names it, in the dossier. That row is designed to be the most persuasive artefact on the page: a system that can be watched saying no to itself, rather than one that asks to be believed.
Model-change control is specified at register scale: 214 models across 1,840 versions, each version carrying a validation state and its unmodelled perils, named. Naming unmodelled perils by name is Lloyd’s own requirement of the market, not our invention — and a model that cannot be registered cannot deploy.
The refused run and the 214-model register are specification, not counts from any live deployment. The governed chain below is what runs today; the specification is what the demo corpus must satisfy before it ships.
The pilot metric
How we count this
Complete means the evaluation re-runs from its seed, the comparison against the incumbent version survives, and a named owner’s sign-off is in the record. A model that could be promoted if somebody re-created the evaluation does not count.The falsifier: a share that rises because the evaluation bar dropped is not readiness. The gate thresholds are reported alongside the share, and a rising share with loosening thresholds is recorded as a failure.The mechanism: a model that cannot be registered cannot deploy
The chain runs preparation, tuning, evaluation gate, registration and deployment as governed steps, each streamed and each recorded — so the evidence for a promotion is produced by the promotion rather than about it.
The evaluation gate is joined by a cross-model check: heterogeneous models re-derive the result and a computed agreement predicate must pass before promotion releases.
A named owner signs at a durable waitpoint, and the signature is part of the model’s record rather than an email attached to it. An unregistrable model has nowhere to go, which is the point.
NX/services/nexus-workflows/src/services/tool-executors/model-{factory,eval,register,deploy}-executor.ts
NCT/services/nexus-workflows/src/services/tool-executors/sentinel-executors.ts
Validation-Evidence Autopilot
The annual validation dossier compiles itself, out of runs that were reproducible on the day they happened — rather than being reconstructed once a year from memory and spreadsheets.
ROS/src/services/gtm-genesis/blueprint/MonteCarloEngine.tsCross-Model Computed Gate → Sealed Certificate
Two or more heterogeneous models re-derive the answer. A computed — not learned — agreement predicate must pass before anything is released, and the inputs and model versions are sealed into the record.
NCT/services/nexus-workflows/src/services/tool-executors/sentinel-executors.tsWhat this path answers
Getting a model approved and keeping it approved are different problems, and this path answers the first one. Whether a model is still the one that was validated is tracked by a person on a calendar, backed by an alert taxonomy and stores that already carry each model’s state. Dispatch tiers 1 and 2 run in production.
REFUSED
When the heterogeneous models disagree at the promotion gate, nothing promotes. The effect is held and the disagreement is named on screen.
The alternative is a promotion that proceeded because one model was confident, which is precisely the event your model-risk committee exists to catch. A held promotion is a bad afternoon; an unheld one is a finding.
What you can do next
Details for support
GATE_DISAGREEMENT
Start with the model that has waited longest
Register one model, run one evaluation through the gate, and have one named owner sign it. What you learn in an afternoon is which parts of your approval record already exist and which parts your team has been re-creating from memory every time.
A person replies with two or three times. Not a sequence.