Skip to content
Decision-Evidence Operating System

Capability deep-dive · deliberation rooms

You added a second opinion to the model. Can you show it read anything the first one did not?

Most “multi-agent” configurations hand every agent the same file, run a few rounds of argument, and take the majority. The published measurement on that configuration is not flattering, and an engineer evaluating a committee of models should start from it rather than from a diagram.

Start with the result that argues against rooms

Two independent author groups report the same finding: vanilla multi-agent debate often underperforms simple majority voting at higher compute cost, and can degrade below a single agent. The two mechanisms they name as missing are diversity of initial viewpoints and explicit, calibrated confidence. Everything below is those two mechanisms turned into refusals, because a documented convention decays and only a mechanism holds.

worsehow vanilla multi-agent debate scored against simple majority voting, at higher cost, in two independent replications
B₂the control a room must beat: n seats, calibration-weighted pool, no argument exchange at all
T1 · T2the dispatch tiers seat invocation and the live room relay run on in production today

The debate result is carried as verified because two independent groups reported consistent findings; the 2026 refinements this design leans on — confidence-weighted voting, minority overturn, log-probability diagnosis of overconfident debaters — are each a single paper, and the room is built so that it does not collapse if they fail to replicate. B₂ is our own baseline ladder, not an external benchmark: B₀ is one seat with the best available bible, B₁ is n seats with an unweighted majority vote, B₂ adds calibration weighting and no argument exchange, and B₃ is the room. A room that cannot separate from B₂ means what you are buying is calibration — and we would rather sell you the calibrated pool and log it than sell you a room.

The seat contract, in four clauses

A seat argues one evidence slice

The evidence admitted for a decision is partitioned into a kernel every seat sees — the notification, the schedule, the loss description — and disjoint slices, one per seat. Overlap above the configured bound does not raise a warning. It refuses the convening, with a machine code, because a room where every seat read the same file is the configuration the measurement above convicts, wearing a room’s costume.

A seat posts a probability, not an opinion

Each seat emits a probability vector — a stated probability for each outcome the room can reach — scored with a proper scoring rule, one that rewards honest probabilities, and broken down into reliability, resolution and uncertainty. A seat that is confident and wrong is measurably distinguishable from a seat that is uncertain and right, and only one of those is a benching signal.

One seat exists to puncture the consensus

The Red-Queen seat does not argue an outcome. It attacks the derivation — the provenance of an artefact, an unstated assumption, a substituted default — and each attack carries a severity and a puncture probability that the survival check consumes before anything reaches the chair.

No seat rules

A seat can conclude. It cannot release. The ruling belongs to a named human chair at a durable waitpoint — a pause that survives any crash of the process, its container or its server and keeps waiting for the person. The effect is idempotent — however many times the machinery retries, it happens at most once — and the ruling event is written before any effect is released. No record, no effect.

rooms, typed messages, persona invocation(shipped)PC/src/routes/rooms.routes.tslive room relay(shipped)NX/services/nexus-gateway/src/websocket/room-event-relay.tsapproval waitpoints(shipped)NX/services/nexus-orchestrator/src/routes/dispatch-routes.tsnever-dropped audit emitter(shipped)NX/services/nexus-workflows/src/services/audit-emitter.ts

A bible is a version, and promoting one is itself a decision

A persona bible is not a prompt in a text box. It is a versioned artefact with declared competence over artefact types, anchored to the clauses it may reason about, carried in the database and in the memory graph together, and evolving through recorded revisions rather than by edit-in-place.

A bible earns a live seat by beating the incumbent on a sequestered drill corpus — practice cases held apart so no seat can train on the test. That corpus is a calibration gym, with leakage rules, a champion–challenger record and an auto-benching rule that stands a seat down when its measured skill decays. Promotion is a governed change that carries its own certificate, so “which version of which persona argued this case” is a question the record answers rather than a question the vendor answers.

evolving bible pattern(shipped)PC/src/blueprint/CharacterBibleManager.tspersona definitions with system prompt and memory config(shipped)PC/migrations/010_persona_definitions.sql

What reaches the chair, and what the dissent map is not allowed to hide

Dissent is computed as a divergence matrix — a grid of how far apart each pair of seats finished — over the seats’ final-round verdicts, weighted by measured calibration — because disagreement from a seat with no measured skill is noise, and a metric that treats it as signal routes noise to the chair. The unweighted spread is reported beside it, so “the disagreement is coming entirely from the weakest seat” is visible rather than hidden inside one index.

Where a calibrated minority’s pooled evidence outweighs the majority’s, the room does not flip the verdict. It raises a flagged minority with the terms of the inequality shown, and puts the objection in front of the accountable human — which is the whole point of having one.

Where the room fails closed

See it in the console

Decision room — in session
Decision room — in sessionSee it in the gallery →

How a case reaches a room

The throughput objection is the right first objection, and it has a design answer rather than a reassurance: the room is where contested decisions go. The routing decision that sends a case to a room instead of down the straight-through lane is itself a logged event, so the proportion of your book that ever sees a committee is a measurable property of your own configuration and not a claim of ours.

What this connects to

Convening triage

What decides that this case needs a room, which chair, and which autonomy tier — and why that routing decision is logged.

Cross-model gates

What the room’s conclusion has to survive before anything is released, and what the agreement predicate does not catch.

Evidence fabric

Where the thread goes: one governed event stream, compiled per regulator rather than reconstructed per audit.

Send me the paper

Four working papers. A person sends the one you pick — no download wall, and no meeting is booked.

Work out what this costs