Deliberation costs more time than one model answering alone. The only honest reason to pay it is that the answer is better often enough to matter. Here is how often, measured against a single frontier model, judged blind — and the same rig, handed to anyone building a mode.
This is the figure the Standard mode’s own listing carries. It is not a marketing number computed for this page — it is read from the same record the product reads.
Every question is classified before it runs. Splitting the result by that classification is the most useful thing we know about deliberation: the gain peaks in the middle and falls off at both ends. On easy questions there is little to deliberate about. On the hardest ones, three models are more likely to be wrong together.
Loading the per-difficulty split…
Measured on the same questions, against the same two single-model baselines. Deliberation is not fast and does not pretend to be.
Loading…
Cost is shown as a multiple rather than a figure — the absolute per-call number is on your own dashboard and your own receipt, where it is yours and it is exact.
A win rate is only worth as much as the protocol behind it. Ours, in full, so it can be argued with.
Our own rig grades our own work on our own question set, which is worth exactly as much as you would expect. So Quorum also ran on a public evaluation platform’s task set, against frontier models we do not control, scored two different ways. To be exact about what this is and is not: we took their tasks and scored the results ourselves. The runs were never submitted to that platform and carry no endorsement from it — the raw log records submitted_to_optima: no on every row. The task set and the opposing models are theirs; the grading is ours.
The two methods disagree about us, sharply, and that is the most useful thing on this page.
20 shared tasks · 6 models · 120 rubric gradings and 300 pairwise comparisons · graded by Amazon’s Nova Pro — a model neither we nor the platform authored, but one we chose and ran. Both axes are truncated to separate the field; neither starts at zero.
Graded against each task’s own correctness rubric, Quorum places second of six — 97.5 against 97.9 for the best frontier model in the set.
Asked which answer a judge simply prefers, head to head, Quorum drops to fourth, on 44.5%. One mid-tier model is far better liked than its correctness score predicts, and we are the mirror image of it.
Being right and being preferred are not the same axis, and we are better at the first. Part of that gap has a known, unglamorous cause we found by running this: the stage that writes the final answer was pinned to an older model than the one taking a seat.
Check it: the derived leaderboard (all six models, both scores, method stated), the rubric gradings, the pairwise comparisons, and the run log.
Models are identified by provider and class. Quorum does not print model identifiers on public pages. Run of 2026-08-18; the underlying judge output is kept in the repository alongside the harness that produced it.
Everything above runs on machinery a Mode Maker can point at their own mode. Iterate while the config is still moving; certify once it has stopped.
A certified score is bound to a config hash. Change the mode after certification and the badge comes off until it is run again — which is the only way a score on a listing means anything.
The record behind the number, published so the win rate is arithmetic rather than a claim. One row per judge verdict — 11,808 of them over 656 rows and 577 distinct prompts, with each judge named, its written reasoning, its tokens, cost and latency.
Download the verdict record (CSV) · JSON · The external benchmark run
Recomputed from that file alone: 62.6% win, 5.8% split, 31.6% loss, against a published 62.5 / 5.9 / 31.6. The one-row gap is 25 verdicts out of 11,808 that returned no winner at all, concentrated on a single question. Model answer bodies are withheld; the full record including them goes to a named investor or researcher on request.