Capability deep-dive · model production
A model reached production last quarter. Can you name its owner, and what it was tested against?
Most model-risk questions are not about the mathematics. They are about whether the chain that put a version behind a decision left anything behind — a dataset with provenance, an evaluation that actually ran, an accountable name, and a way back. A dashboard does not answer that. A refusal does.
Five steps, one job, one causal order
step 1
prep
prepared dataset — dataset lineage required before anything reads it
step 2
tune
checkpoint — submitted as a batch job through one gateway; progress streams back live, not polled blind
step 3
eval-gate
eval report — the only step that can stop the promotion
step 4
register
registered model — a named owner, or no registration
step 5
deploy
serving endpoint — only a model that actually registered
There is no pipeline definition file. The chain is one row in the skill registry — the platform’s catalogue of governed capabilities — and the five steps run as a single job with one identity, one progress stream and one causal order — so the thread for a training run is an unbroken chain rather than five loosely correlated jobs someone has to correlate afterwards. Each step’s output is stored under its own named key and handed to the next step along an explicitly declared path, so what depends on what can be read from the record rather than inferred from the code.
Counts are read from our own migrations, executors and repositories. Three of the six alert types — latency, error rate and cost — are computable from telemetry the plane already holds. The provider count is a registration count, not an availability claim.
A model may argue inside a governed chain. It may never be how a step is invoked
This rule was not derived from principle. It was measured. Every step originally invoked its tool through a reasoning loop, and the model did not reliably emit the tool call — even with the call forced — so the executor never ran, the step output was an empty string, and the downstream steps fell into a skip cascade that read to an operator as “the evaluation report is empty”.
A LIVENESS FAILURE PRESENTING AS A CONTENT FAILURE
That is the general shape, and it is why the invocation seam is deterministic now. A step that is invoked probabilistically can fail by never running at all — a liveness failure — and that failure arrives looking like a content problem. Under log-or-refuse — no record, no effect — an unexecuted step must be a refusal event with a code — and a probabilistic invoker cannot guarantee it produces one, because it may also fail to emit the refusal. Determinism at the invocation seam is what makes the run ledger’s silence mean something.
A second measured defect is worth carrying for the same reason: a queue name off by one value produced a dispatch that resolved, enqueued and returned success onto a queue no worker consumed. The run sat queued forever while everything else processed normally. A dispatch acknowledgement is not evidence that work happened. Only the run ledger is.
The promotion gate
Registration runs only if the evaluation shows no quality regression and no safety regression, and deployment runs only if registration actually happened. That gate is expressed in the step definition and enforced in the executor, which stops the promotion before the registry is even consulted — defence in depth that exists because the team hit the invocation defect above and did not want the gate to depend on a prompt.
AN UNREGISTRABLE MODEL CANNOT DEPLOY
No named owner, no registration. No registration, no endpoint. The chain is the control rather than a checklist somebody signs, which is a mechanism a model-risk reviewer can test in an afternoon. The assurance rung a version reaches — unevaluated, evaluated, cross-checked, board-approved, refused — decides what it is allowed to do, and a model may be cross-checked and still not touch a filed rating chain.
the five step executors(shipped)NX/services/nexus-workflows/src/services/tool-executorsthe single accelerated-compute path(shipped)NX/services/nexus-hpc-gatewaythe model plane of record(shipped)ROS/src/routes/gpu-ml.ts
Where the chain fails closed
REFUSED
A promotion to a filed decision class was attempted without a sealed ruling from the named owner. The version was not registered for that class and no endpoint was cut over.
Promotion is a change to decision behaviour, so it is a governed decision in its own right rather than a deploy button — proposal, deliberation, cross-model gate, ruling with mandatory reasons, and a rollback pointer to the previous serving version, all on one thread. That is what makes a regulator’s treatment of a new model version as a capital event mechanisable instead of aspirational. A deploy step will not serve a model whose registration status is anything but registered, and a registration will not proceed on a reported regression.
Details for support
AUTHORITY_EXCEEDED
What this connects to
Validation evidence
What the run ledger is for: a dossier compiled out of runs that were reproducible on the day they happened.
Stress testing
The other place run identity decides admissibility — and where a back-test cannot validate what it looks like it validates.
Cross-model gates
The rung between evaluated and board-approved, applied to a model’s own outputs on a sequestered fold — held-out test data the model never saw.
Four working papers. A person sends the one you pick — no download wall, and no meeting is booked.
Work out what this costs