Migrating from OpenAI

Change the base URL and set the model to a Mode ticker. What carries over from the OpenAI chat completions API and what Quorum adds to the response.

Quorum serves the chat completions shape at POST /v1/chat/completions. Existing OpenAI client code runs against it after two changes: the base URL, and the model value.

The two changes

from openai import OpenAI

# Base URL: was https://api.openai.com/v1. API key: a Quorum key.
# Timeout: a deliberation runs several models.
client = OpenAI(
    base_url="https://www.quorum.dog/v1",
    api_key="qk_live_...",
    timeout=300,
)

# Model: was "gpt-4o", now a Mode ticker.
r = client.chat.completions.create(
    model="QRUM:STAN",
    messages=[{"role": "user", "content": "Should we shard this table now or after launch?"}],
)
print(r.choices[0].message.content)
import OpenAI from 'openai';

const client = new OpenAI({
  baseURL: 'https://www.quorum.dog/v1',
  apiKey: process.env.QUORUM_API_KEY,
  timeout: 300_000,
});

const r = await client.chat.completions.create({
  model: 'QRUM:STAN',
  messages: [{ role: 'user', content: 'Should we shard this table now or after launch?' }],
});
console.log(r.choices[0].message.content);

A ticker has the form COMPANY:MODE. model names a Mode, and a Mode decides which engines sit in the seats and how the panel runs. For example, QRUM:STAN is Standard Mode. GET /v1/models lists the Modes your key can call, with pricing.

What carries over

OpenAIQuorum
POST /v1/chat/completionsSame path and method.
Authorization: BearerSame header, with a Quorum key (qk_live_... or qk_test_...).
messages with role and contentSame. Send content as a string.
id, object, created, modelSame fields. object is chat.completion.
choices[0].message.contentThe answer.
choices[0].finish_reasonstop, or length when the answer was cut off.
usage.prompt_tokens, completion_tokens, total_tokensSame fields.
usage.completion_tokens_details.reasoning_tokensSame field. null when the provider reported no breakdown.
GET /v1/modelsSame path. Returns the Modes your key can call.
Error body{ "error": { "message", "type", "code", "param" } }, the OpenAI shape.
Idempotency-Key headerSame header. A repeat of a call that succeeded replays the answer with quorum.replayed: true. The replay matches on the key alone, so send a new key for each request. A call that ended capped is not replayed, so a retry runs and bills again. If you omit the header, Quorum derives a key from your organization, key, model and messages, so an SDK retry of the same request replays instead of billing twice.

The request fields Quorum reads are model, messages, tier, quorum.max_cost_usd and stream. A Mode's own configuration sets how its panel runs.

What Quorum adds

The response carries a quorum block next to the standard fields. Read it to see how the seats came out.

{
  "id": "…",
  "object": "chat.completion",
  "choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
  "usage": { "prompt_tokens": 41, "completion_tokens": 612, "total_tokens": 653 },
  "quorum": {
    "request_id": "…",
    "mode": "QRUM:STAN",
    "rounds": 1,
    "seats": 3,
    "seats_answered": 3,
    "converged": true,
    "remaining_friction": null,
    "contested_passage": null,
    "capped": false,
    "cap_credit_usd": 0,
    "billed_usd": 0.10,
    "latency_ms": 26418,
    "engines": ["<seat model id>", "<seat model id>", "<seat model id>"]
  }
}

The values in the sample are illustrative. A replay of an earlier call carries a shorter quorum block with replayed: true: no converged, remaining_friction, contested_passage or billed_usd. Check replayed first.

FieldMeaning
request_idThe id for GET /v1/receipts/{request_id}.
convergedfalse when the seats did not settle on one position. null when no judge read the round, as on the express lane.
remaining_frictionfactual_conflict when the judge flagged a factual conflict between seats, otherwise null.
contested_passageAn object { seat, quote, why }, or null when no single passage carried the split.
seats, seats_answeredSeats convened and seats that answered.
billed_usd, surcharge_usd, platform_fee_usd, cap_credit_usdThe bill for this call and its parts. billed_usd is the token charge, the platform fee and the surcharge, minus cap_credit_usd.
cappedtrue when the call went over its ceiling. The call is charged in full and the amount over the ceiling is credited back in cap_credit_usd. The answer is whole.
enginesThe model ids of the seats that ran.
substitutionsSeats that ran a backup engine, and what ran instead.

Settings that differ

Migrating an agent

An agent that called OpenAI on every step can keep that model for routine steps and call Quorum at decision points. The escalation router shows how to choose per question, and Agent frameworks wraps the call as a tool.

Next