Skip to content
Decision-Evidence Operating System

Capability deep-dive · cross-model gates

Your models agreed. What does that actually rule out?

A second model that confirms the first is worth something only if you can say what it read, how far apart the two answers were allowed to be, and how correlated their errors are. Without those three numbers, “two models checked it” is unfalsifiable.

The predicate is computed, not learned

Two or more heterogeneous models re-derive the answer, and a predicate — a fixed yes-or-no rule — over their structured outputs must pass before anything is released. It is arithmetic, not a judge model, so its inputs are inspectable and its outcome reproduces. It has three conjuncts — three conditions that must all hold at once — and dropping any one of them is how a gate becomes theatre.

conjunct 1

Categorical identity

Every categorical field of the structured output must be identical across the models. Not similar, not semantically close — identical.

conjunct 2

Numeric tolerance, per field

The spread between the highest and lowest value must sit inside a tolerance declared for that field and that decision class. Tightening the tolerance raises refusals; it does not lower escapes — the wrong answers that pass.

conjunct 3

Evidence overlap

The models’ cited evidence sets must overlap above a floor. Two models that reach the same answer while reading different documents do not pass — agreement on a conclusion is not agreement on a derivation.

35 of 102cross-model validators proven live in the deployed core, drawn from our own capability record
≥ 2heterogeneous models that must re-derive before the predicate is evaluated at all
φthe measured error correlation of the model pair, named on the certificate rather than assumed away

The predicate, the verdict and the record are built.

“35 of 102” is a count of validators observed live in the deployed cross-model validation core, carried verbatim from our own capability record — not a coverage percentage and not a quality score. It is a governance input rather than trivia: a warning class whose validator is not among those live cannot be gated, so it is delivered as an explicitly ungated observation and may not propose an action. The pair correlation is estimated on a shared benchmark of decided cases; where a pair exceeds its admissibility bound — the maximum error correlation allowed — the gate is refused rather than reported.

What the gate does not catch

It does not decide the question

The gate decides what a claim is permitted to say. A passing predicate yields a gated warning that may propose an action. A failing one yields a contested warning where the disagreement is the content, and it may not propose anything. A class with no live validator yields an ungated observation, labelled as such. A dead input yields a refusal and no warning at all.

It never deletes the disagreement

A failed gate does not suppress the warning. Deletion is the failure mode a governance layer exists to prevent: it produces a system that looks calm because its disagreements are invisible. The contested class is how a disagreement stays visible, with both derivations openable side by side.

A gate is capped by its worst input, and can never raise a tier

Every feed datum has three clocks, not one: the time the observation describes, the time we received it, and the time a decision read it. The difference between the first and third governs truth; the other two assign blame correctly, which matters when an operational-resilience report has to say whether a degradation was ours.

Staleness composes downward and only downward. Adding an input can never improve a decision’s freshness, so a pipeline cannot launder a dead feed by averaging it with a live one; unknown provenance resolves to dead rather than to fresh; and a decision that consumed a stale input carries that fact permanently, because it is sealed with the thread. When a feed degrades, the achievable assurance tier drops and the certificate says so. It does not substitute a default.

cross-model validation core(shipped)NCT/services/nexus-workflows/src/services/tool-executors/sentinel-executors.tsno silent fallback on a missing upstream key(shipped)SOV/services/nexus-weather-feeds/src/lib.rs

A watcher may propose. It may not arm

A watcher that reads a hurricane advisory or a news stream can propose an effect — open an event room, review a moratorium, re-price an accumulation. The transition from proposed to armed requires a named human through a gated binding, and the armed state shows who armed it. Irreversibility overrides value: a signal that can trigger a spent effect — one that cannot be taken back once released — gates against a third model regardless of how cheap it looked.

The second law is smaller and matters as much: a dead feed is a loud row, never a missing row. A watcher list that only ever shows healthy sources has told you nothing about what it does when a source dies.

Where the gate fails closed

What this connects to

The formal lane

Where a question is decidable, the answer can carry a checkable witness instead of a convergence. The gate is what covers everywhere it does not reach.

Deliberation rooms

What produces the conclusion the gate then re-derives — and the chair whose ruling the gate sits in front of.

Trust centre

The standing register behind every claim on this page, with dates.

Send me the paper

Four working papers. A person sends the one you pick — no download wall, and no meeting is booked.

Work out what this costs