DEV · local workspace

Skip to content

Reference data across observed endpoints

A market price for inference capacity

Megatron Markets provides market infrastructure for spot and forward-dated AI-model capacity across providers and models. The product compares standardised bundles of prompt, cache-read, cache-write and output tokens across clearly identified terms.

Z.ai GLM 5.2$0.4886$2.314.7×DeepSeek V4 Flash 0731$0.0786$0.44005.6×MoonshotAI Kimi K2.6$0.5605$1.202.1×OpenAI gpt-oss-120b$0.0300$0.350011.7×DeepSeek V4 Flash 0423$0.0679$0.44006.5×Z.ai GLM 5.1$0.8960$1.541.7×Google Gemma 4 31B$0.0800$0.990012.4×DeepSeek V4 Pro 0423$0.6600$1.912.9×MoonshotAI Kimi K2.7 Code$0.6700$1.902.8×DeepSeek V3.2$0.2088$3.0014.4×MoonshotAI Kimi K3$2.60$6.002.3×Meta Llama 3.3 70B Instruct$0.1000$1.0410.4×
414
Models observed
102
Inference providers observed
1,505
Provider endpoints observed
28
Providers on the most observed model

Reference observation of provider-published rates. Read at 17 August 2026, 17:19 UTC.

Mechanics

From reference data to managed workflows

01

Compare

Provider-published rates and declared capacity are organized as reference data. A desk can inspect the standard bundle shape and compare the inputs.

02

Configure

The platform supports a central limit order book for spot and forward-dated instruments when the selected instrument is enabled for new exposure, with account-scoped controls and attributed decisions.

03

Operate

Order state, fills, cancellations and desk enquiries connect into one product workflow with records for positions, fees and capacity evidence.

Reference data

Provider rate dispersion

A model is often represented by several provider-published rate cards, so a reference price is a range rather than a single number. Each bar spans the cheapest and the dearest prompt rate observed in the dated reference set.

203
models have reference observations from more than one provider.
179
of those show more than one published rate, so the difference is not a rounding artefact.
2.2×
is the median ratio of dearest to cheapest across that group; at the ninetieth percentile it is 4.4×.

Prompt rate in US dollars per million tokens, cheapest to dearest, on a logarithmic axis. The figure beside each model is its count of observed endpoints. Reference observation of provider-published rates. Read at 17 August 2026, 17:19 UTC.

Reference

The reference catalogue

Models and provider-published rates, medians and spreads in one view. This dated reference observation shows where rates differ before a desk evaluates an instrument or builds a bundle.

The Megatron Markets reference catalogue: a table of models with author, context length, provider count, prompt and output rates, median rate and spread in basis points.
Reference observation · provider rates and endpoints, read 17 August 2026

Instruments

The instrument unit

The bundle builder, showing a split across prompt, cache and output-token legs with a reference value beneath it.

Bundle

A workload-shaped bundle

The instrument is a basket naming a weight per leg — prompt, cache-read, cache-write and output — with weights summing to one and a reference value derived from selected leg rates.

A model card showing a table of provider endpoints with prompt, output and cache prices.

Matrix

The provider rate matrix

Every model card carries one row per observed provider endpoint: published prompt, output and cache rates, the quantisation in use, and the distance from the cheapest observed rate in basis points.

Forward design

Tenor and delivery window

The tenor states how far ahead the named window opens; the window states the future period represented by the instrument, while the window defines the period being priced.

Supply

Declared capacity and its floor

A provider can describe capacity, a price floor and a validity window. Those fields travel with the provider and model identity used for comparison.

Access

Book, RFQ and desk enquiry

The product supports a central order book, requests for quote and desk enquiries, with each path preserving participant and instrument identity.

Context

How inference capacity is bought today

Published rates make comparison possible. Standardised instruments add model, provider, token leg and time-window detail that a single headline rate cannot carry.

01

Price comparison

Published rate cards provide a dated reference point. When an active fixed-window forward product is available, it adds time detail without requiring every workload to share one commitment shape.

02

Provider comparison

The same model can carry materially different rates across endpoints. A common reference catalogue makes that dispersion visible before a desk decides what to investigate.

03

Idle capacity

Idle accelerator time is perishable. For an active product, provider curve points express available MFT, minimum clip, price and firm expiry for a named model and delivery window.

04

The shape of demand

A short delivery window and a long one impose different capacity constraints. Naming tenor and window makes time part of a planned instrument rather than an unstated assumption.

Precedent

A market-design analogy, not the product

Electricity-market research is useful because it studies capital-intensive, perishable capacity under changing demand. It is a structural comparison only. The product here is AI-model inference capacity, represented by token bundles across providers and models.

01

Capital-intensive supply

Both electricity generation and AI inference depend on expensive physical capacity. That makes availability and utilisation economically important in each market.

02

Supply that cannot be stored

Idle accelerator time cannot be placed into inventory for tomorrow. That perishability is the useful comparison with electricity, even though the delivered products and operational constraints differ.

03

Time-shaped markets

Electricity research distinguishes near-delivery and future markets. The transferable idea is to name time explicitly: Megatron Markets is designing spot and forward-dated AI-capacity instruments with defined windows.

04

Distinct product identity

The product uses its own provider, model, token-leg and time-window identity. Electricity-market research informs comparison questions without changing those product fields.