Capability deep-dive · deliberation rooms
You added a second opinion to the model. Can you show it read anything the first one did not?
Most “multi-agent” configurations hand every agent the same file, run a few rounds of argument, and take the majority. The published measurement on that configuration is not flattering, and an engineer evaluating a committee of models should start from it rather than from a diagram.
Start with the result that argues against rooms
Two independent author groups report the same finding: vanilla multi-agent debate often underperforms simple majority voting at higher compute cost, and can degrade below a single agent. The two mechanisms they name as missing are diversity of initial viewpoints and explicit, calibrated confidence. Everything below is those two mechanisms turned into refusals, because a documented convention decays and only a mechanism holds.
The debate result is carried as verified because two independent groups reported consistent findings; the 2026 refinements this design leans on — confidence-weighted voting, minority overturn, log-probability diagnosis of overconfident debaters — are each a single paper, and the room is built so that it does not collapse if they fail to replicate. B₂ is our own baseline ladder, not an external benchmark: B₀ is one seat with the best available bible, B₁ is n seats with an unweighted majority vote, B₂ adds calibration weighting and no argument exchange, and B₃ is the room. A room that cannot separate from B₂ means what you are buying is calibration — and we would rather sell you the calibrated pool and log it than sell you a room.
The seat contract, in four clauses
A seat argues one evidence slice
The evidence admitted for a decision is partitioned into a kernel every seat sees — the notification, the schedule, the loss description — and disjoint slices, one per seat. Overlap above the configured bound does not raise a warning. It refuses the convening, with a machine code, because a room where every seat read the same file is the configuration the measurement above convicts, wearing a room’s costume.
A seat posts a probability, not an opinion
Each seat emits a probability vector — a stated probability for each outcome the room can reach — scored with a proper scoring rule, one that rewards honest probabilities, and broken down into reliability, resolution and uncertainty. A seat that is confident and wrong is measurably distinguishable from a seat that is uncertain and right, and only one of those is a benching signal.
One seat exists to puncture the consensus
The Red-Queen seat does not argue an outcome. It attacks the derivation — the provenance of an artefact, an unstated assumption, a substituted default — and each attack carries a severity and a puncture probability that the survival check consumes before anything reaches the chair.
No seat rules
A seat can conclude. It cannot release. The ruling belongs to a named human chair at a durable waitpoint — a pause that survives any crash of the process, its container or its server and keeps waiting for the person. The effect is idempotent — however many times the machinery retries, it happens at most once — and the ruling event is written before any effect is released. No record, no effect.
rooms, typed messages, persona invocation(shipped)PC/src/routes/rooms.routes.tslive room relay(shipped)NX/services/nexus-gateway/src/websocket/room-event-relay.tsapproval waitpoints(shipped)NX/services/nexus-orchestrator/src/routes/dispatch-routes.tsnever-dropped audit emitter(shipped)NX/services/nexus-workflows/src/services/audit-emitter.ts
A bible is a version, and promoting one is itself a decision
A persona bible is not a prompt in a text box. It is a versioned artefact with declared competence over artefact types, anchored to the clauses it may reason about, carried in the database and in the memory graph together, and evolving through recorded revisions rather than by edit-in-place.
A bible earns a live seat by beating the incumbent on a sequestered drill corpus — practice cases held apart so no seat can train on the test. That corpus is a calibration gym, with leakage rules, a champion–challenger record and an auto-benching rule that stands a seat down when its measured skill decays. Promotion is a governed change that carries its own certificate, so “which version of which persona argued this case” is a question the record answers rather than a question the vendor answers.
evolving bible pattern(shipped)PC/src/blueprint/CharacterBibleManager.tspersona definitions with system prompt and memory config(shipped)PC/migrations/010_persona_definitions.sql
What reaches the chair, and what the dissent map is not allowed to hide
Dissent is computed as a divergence matrix — a grid of how far apart each pair of seats finished — over the seats’ final-round verdicts, weighted by measured calibration — because disagreement from a seat with no measured skill is noise, and a metric that treats it as signal routes noise to the chair. The unweighted spread is reported beside it, so “the disagreement is coming entirely from the weakest seat” is visible rather than hidden inside one index.
THE MAP DEGRADES BEFORE IT LIES
The dissent map is a two-dimensional flattening of a disagreement that has more dimensions than two, so it carries its stress statistic — the measure of how much the flattening distorted — on its face. Above the stress limit the map is replaced by the raw divergence matrix rather than drawn. A dissent map that always looks clean is a dissent map that is lying about its own geometry.
Where a calibrated minority’s pooled evidence outweighs the majority’s, the room does not flip the verdict. It raises a flagged minority with the terms of the inequality shown, and puts the objection in front of the accountable human — which is the whole point of having one.
Where the room fails closed
REFUSED
The chair’s waitpoint reached its deadline without a ruling. The room produced no verdict, and no effect was released.
A deadlocked room with a chair timeout has an honest terminal state, and it is not “proceed”. Deadline expiry escalates up the authority matrix as a new waitpoint; it never releases the decision on its own. The same law makes the reverse true: the ruling event is written durably before the effect happens, so a ruling that could not be recorded is a ruling that did not take effect.
What you can do next
NX/services/nexus-orchestrator/src/routes/dispatch-routes.tsDetails for support
DEADLINE_EXCEEDED
See it in the console

How a case reaches a room
The throughput objection is the right first objection, and it has a design answer rather than a reassurance: the room is where contested decisions go. The routing decision that sends a case to a room instead of down the straight-through lane is itself a logged event, so the proportion of your book that ever sees a committee is a measurable property of your own configuration and not a claim of ours.
What this connects to
Convening triage
What decides that this case needs a room, which chair, and which autonomy tier — and why that routing decision is logged.
Cross-model gates
What the room’s conclusion has to survive before anything is released, and what the agreement predicate does not catch.
Evidence fabric
Where the thread goes: one governed event stream, compiled per regulator rather than reconstructed per audit.
Four working papers. A person sends the one you pick — no download wall, and no meeting is booked.
Work out what this costs