- 414
- Models observed
- 102
- Inference providers observed
- 1,505
- Provider endpoints observed
- 28
- Providers on the most observed model
Reference observation of provider-published rates. Read at 17 August 2026, 17:19 UTC.
Mechanics
From reference data to managed workflows
Compare
Provider-published rates and declared capacity are organized as reference data. A desk can inspect the standard bundle shape and compare the inputs.
Configure
The platform supports a central limit order book for spot and forward-dated instruments when the selected instrument is enabled for new exposure, with account-scoped controls and attributed decisions.
Operate
Order state, fills, cancellations and desk enquiries connect into one product workflow with records for positions, fees and capacity evidence.
Reference data
Provider rate dispersion
A model is often represented by several provider-published rate cards, so a reference price is a range rather than a single number. Each bar spans the cheapest and the dearest prompt rate observed in the dated reference set.
- 203
- models have reference observations from more than one provider.
- 179
- of those show more than one published rate, so the difference is not a rounding artefact.
- 2.2×
- is the median ratio of dearest to cheapest across that group; at the ninetieth percentile it is 4.4×.
Prompt rate in US dollars per million tokens, cheapest to dearest, on a logarithmic axis. The figure beside each model is its count of observed endpoints. Reference observation of provider-published rates. Read at 17 August 2026, 17:19 UTC.
Reference
The reference catalogue
Models and provider-published rates, medians and spreads in one view. This dated reference observation shows where rates differ before a desk evaluates an instrument or builds a bundle.


Instruments
The instrument unit

Bundle
A workload-shaped bundle
The instrument is a basket naming a weight per leg — prompt, cache-read, cache-write and output — with weights summing to one and a reference value derived from selected leg rates.

Matrix
The provider rate matrix
Every model card carries one row per observed provider endpoint: published prompt, output and cache rates, the quantisation in use, and the distance from the cheapest observed rate in basis points.
Forward design
Tenor and delivery window
The tenor states how far ahead the named window opens; the window states the future period represented by the instrument, while the window defines the period being priced.
Supply
Declared capacity and its floor
A provider can describe capacity, a price floor and a validity window. Those fields travel with the provider and model identity used for comparison.
Access
Book, RFQ and desk enquiry
The product supports a central order book, requests for quote and desk enquiries, with each path preserving participant and instrument identity.
Context
How inference capacity is bought today
Published rates make comparison possible. Standardised instruments add model, provider, token leg and time-window detail that a single headline rate cannot carry.
Price comparison
Published rate cards provide a dated reference point. When an active fixed-window forward product is available, it adds time detail without requiring every workload to share one commitment shape.
Provider comparison
The same model can carry materially different rates across endpoints. A common reference catalogue makes that dispersion visible before a desk decides what to investigate.
Idle capacity
Idle accelerator time is perishable. For an active product, provider curve points express available MFT, minimum clip, price and firm expiry for a named model and delivery window.
The shape of demand
A short delivery window and a long one impose different capacity constraints. Naming tenor and window makes time part of a planned instrument rather than an unstated assumption.
Precedent
A market-design analogy, not the product
Electricity-market research is useful because it studies capital-intensive, perishable capacity under changing demand. It is a structural comparison only. The product here is AI-model inference capacity, represented by token bundles across providers and models.
Capital-intensive supply
Both electricity generation and AI inference depend on expensive physical capacity. That makes availability and utilisation economically important in each market.
Supply that cannot be stored
Idle accelerator time cannot be placed into inventory for tomorrow. That perishability is the useful comparison with electricity, even though the delivered products and operational constraints differ.
Time-shaped markets
Electricity research distinguishes near-delivery and future markets. The transferable idea is to name time explicitly: Megatron Markets is designing spot and forward-dated AI-capacity instruments with defined windows.
Distinct product identity
The product uses its own provider, model, token-leg and time-window identity. Electricity-market research informs comparison questions without changing those product fields.

