More engines, more sources

An engine is a model. A source is a place that serves it. Quorum picks the cheapest healthy source for each call.

An engine is a model. A source is a place that serves it: a provider, an endpoint and a price. One engine can have several sources, and adding a source gives a model another way to be served.

A Mode names the engines for its seats. Quorum decides which source serves each one.

How a source is picked

For each seat, Quorum considers the sources that carry the Mode's engine and can take traffic.

  1. Health. Sources are held to a health floor: 0.70, relaxed to 0.50 when no source clears it, then any source.
  2. Price. The remaining sources are ranked by what the call would cost the caller: the source's real cost for a typical call, plus its platform fee. The lowest wins.
  3. Latency. A tie goes to the faster source.

Real cost includes the provider's own markup and per-request charge, so a source that looks cheap per token but adds a fee is ranked on the total.

A source with no price, or a price of zero, ranks below every priced source.

When a source fails

If the chosen source fails, Quorum tries the other sources of the same engine first. After a timeout or a provider-wide error, sources on other providers go first, since the same provider would likely fail again. If no source for that engine can run, the seat moves to a backup the Mode approved: the backup it set for that seat, or another engine from the pool a dynamic Mode approved. When a Mode is resolved, the seat stays inside its approved engines.

If neither the engine nor its backup can run, the seat is refused and the receipt says why: unavailable: reason. See Reading a receipt.

What a new model needs

A new model enters the catalog with an exact published price and a source that passes a probe. Once live, the Mode builder can seat it, and a Mode that draws from a pool can pick it up. A Mode that names its engines keeps them until its maker changes them.

Staying current

A probe job is scheduled every 15 minutes. It re-tests the lowest-scoring and longest-unchecked sources, so a recovered source returns to ranking and a failing one drops out. A sync scheduled once a week checks each provider for new and retired models.