Testing

We publish the score
we would rather not.

Deliberation costs more time than one model answering alone. The only honest reason to pay it is that the answer is better often enough to matter. Here is how often, measured against a single frontier model, judged blind — and the same rig, handed to anyone building a mode.

01 — Standard, measured

The number

Preferred over a single frontier model
63.9% ±3.9
577 questions, 4 batches, blind pairwise judging. 3.7% came back split. 32.4% were losses — a single model was preferred. Win, split and loss are question-wise and sum to 100. The row-wise loss rate is 31.6% against a different denominator; Method §07 explains the gap.
Published 2026-08-11. Live figure loads with the page.

This is the figure the Standard mode’s own listing carries. It is not a marketing number computed for this page — it is read from the same record the product reads.

02 — Where the gain is

By question difficulty

Every question is classified before it runs. Splitting the result by that classification is the most useful thing we know about deliberation: the gain peaks in the middle and falls off at both ends. On easy questions there is little to deliberate about. On the hardest ones, three models are more likely to be wrong together.

Loading the per-difficulty split…

03 — The other side of it

What the win costs

Measured on the same questions, against the same two single-model baselines. Deliberation is not fast and does not pretend to be.

Loading…

Cost is shown as a multiple rather than a figure — the absolute per-call number is on your own dashboard and your own receipt, where it is yours and it is exact.

04 — Method

How the score is produced

A win rate is only worth as much as the protocol behind it. Ours, in full, so it can be argued with.

01
Real questions, not ours
Prompts are drawn from a public conversation dataset — things people actually asked, not a suite we wrote to look good on.
02
The baseline is a real frontier model
Each question also runs against two single frontier models, through their own APIs, at the same moment. Not a weakened control.
03
Judges are blind and the order is shuffled
Judges see two anonymous answers, A and B, in randomised order. Nothing identifies which came from deliberation.
04
Three judges, from three different providers
A single judge would import a single lab’s taste. Each question collects verdicts from three independent judges against each of the two baselines, repeated, so a lucky pass does not carry a question.
05
A question is only won on a majority
Ties and even splits are reported as splits, not distributed to whoever needs them.
06
The interval is published with the number
Any figure quoted without its interval and its n should be treated as marketing, including ours. Both are above.
07
And the two aggregations do not agree
The published figure counts a question once, averaging its two matchups, and treats a judge panel that could not decide as a loss. The recomputation counts a row, pools every judge verdict from both matchups into one tally, and reports an undecided panel as its own third state. Both are defensible; they differ in the unit of analysis and in what to do with a tie. We show the gap rather than pick the higher one.
05 — Graded against models we do not control

On someone else’s tasks

Our own rig grades our own work on our own question set, which is worth exactly as much as you would expect. So Quorum also ran on a public evaluation platform’s task set, against frontier models we do not control, scored two different ways. To be exact about what this is and is not: we took their tasks and scored the results ourselves. The runs were never submitted to that platform and carry no endorsement from it — the raw log records submitted_to_optima: no on every row. The task set and the opposing models are theirs; the grading is ours.

The two methods disagree about us, sharply, and that is the most useful thing on this page.

20 shared tasks · 6 models · 120 rubric gradings and 300 pairwise comparisons · graded by Amazon’s Nova Pro — a model neither we nor the platform authored, but one we chose and ran. Both axes are truncated to separate the field; neither starts at zero.

What it says

Graded against each task’s own correctness rubric, Quorum places second of six — 97.5 against 97.9 for the best frontier model in the set.

Asked which answer a judge simply prefers, head to head, Quorum drops to fourth, on 44.5%. One mid-tier model is far better liked than its correctness score predicts, and we are the mirror image of it.

Being right and being preferred are not the same axis, and we are better at the first. Part of that gap has a known, unglamorous cause we found by running this: the stage that writes the final answer was pinned to an older model than the one taking a seat.

Check it: the derived leaderboard (all six models, both scores, method stated), the rubric gradings, the pairwise comparisons, and the run log.

Models are identified by provider and class. Quorum does not print model identifiers on public pages. Run of 2026-08-18; the underlying judge output is kept in the repository alongside the harness that produced it.

06 — The same rig, for your mode

Test yours the way we test ours

Everything above runs on machinery a Mode Maker can point at their own mode. Iterate while the config is still moving; certify once it has stopped.

Step 01
Test
Run your mode against a question set and get per-question verdicts back. The config stays editable — change a seat, run it again, compare.
25% off the mode’s own price · dry runs are free
Step 02
Certify
The standard set and the full protocol: 45 deliberations spread evenly across light, medium and hard, three judges, two repeats, against a locked config hash. This is the run that produces the score, and it is the one that costs.
Billed at real compute cost, capped · about forty-five minutes
Step 03
List
A certified mode carries its score into the Marketplace, next to ingredients derived from its real configuration rather than its description.
Revenue share on use

A certified score is bound to a config hash. Change the mode after certification and the badge comes off until it is run again — which is the only way a score on a listing means anything.

Build and test a mode See what’s published

Check it yourself

The record behind the number, published so the win rate is arithmetic rather than a claim. One row per judge verdict — 11,808 of them over 656 rows and 577 distinct prompts, with each judge named, its written reasoning, its tokens, cost and latency.

Download the verdict record (CSV)  ·  JSON  ·  The external benchmark run

Recomputed from that file alone: 62.6% win, 5.8% split, 31.6% loss, against a published 62.5 / 5.9 / 31.6. The one-row gap is 25 verdicts out of 11,808 that returned no winner at all, concentrated on a single question. Model answer bodies are withheld; the full record including them goes to a named investor or researcher on request.