As of 23 August 2026, the cheapest published output price among major LLM APIs sits at $0.40 per million output tokens for Google's Gemini 2.5 Flash-Lite, while Anthropic's flagship Claude Fable 5 is listed at $50.00 per million output tokens — two entries in the same live pricing landscape. Between those extremes sit Anthropic's Claude Haiku 4.5, Sonnet 5 and Opus 5, Google's Gemini 3.1 Pro, and OpenAI's gpt-5.4-nano, each dated to the comparison page that published it, because vendors increasingly reprice, add context-length tiers, or run time-limited introductory rates without a formal changelog. This tracker compiles every verifiable per-token figure published between April and August 2026, dates each one to its source, and works through what the numbers mean for UK IT buyers weighing build-vs-buy decisions.
View the data behind this chart
| Input $/1M… | Output $/1M… | Source date | |
|---|---|---|---|
| gpt-5.4-nano | $0.20 | $1.25 | Jul 2026 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | Jul 2026 |
| Claude Haiku 4.5 | $1.00 | $5.00 | Aug 2026 |
| Gemini 3.1 Pro (≤200K) | $2.00 | $12.00 | Aug 2026 |
| Claude Sonnet 5… | $2.00 | $10.00 | Aug 2026 |
| Claude Opus 5 | $5.00 | $25.00 | Aug 2026 |
| Claude Fable 5 | $10.00 | $50.00 | Aug 2026 |
August 2026 Snapshot: What Actually Changed
CloudZero's August 2026 comparison lists Anthropic's current line-up as Claude Haiku 4.5 at $1.00 per million input tokens and $5.00 per million output tokens, Claude Sonnet 5 at $2.00/$10.00, Claude Opus 5 at $5.00/$25.00, and the top-tier Claude Fable 5 at $10.00/$50.00 per million tokens.
Separately, Spheron's August 2026 tracker put Google's Gemini 3.1 Pro at $2.00 input and $12.00 output per million tokens for prompts at or below 200,000 tokens, rising to $4.00/$18.00 above that threshold — and flagged that Claude Sonnet 5's pricing is an introductory rate running through 31 August 2026, with a separate standard rate scheduled to begin 1 September 2026.
At the budget end, VibeEngines' July 2026 handbook recorded Gemini 2.5 Flash-Lite at $0.10 input / $0.40 output per million tokens — the cheapest entry-tier figure in this dataset — alongside OpenAI's gpt-5.4-nano at $0.20/$1.25 and Anthropic's Claude Haiku 4.5 at $1.00/$5.00, the same Haiku figure independently confirmed a month later by CloudZero's August snapshot.

How Token Pricing Actually Works: Tiers, Thresholds and Quiet Repricing
Every provider in this dataset charges output tokens at a materially higher rate than input tokens — a pattern worth checking before estimating cost from input volume alone. Readers new to the mechanics may find it useful to understand LLM tokens and context windows before comparing rate cards, since 'a million tokens' is not the same unit of work across every model or task.
Context length is now part of the price, not just a spec. SaaSTweaks' May 2026 comparison shows Google's Gemini 2.5 Pro billed at $1.25 per million input tokens for prompts under 200,000 tokens, rising to $2.50 above that line. Spheron's August 2026 figures show the same pattern for Gemini 3.1 Pro, where both input and output rates roughly double once a prompt crosses the 200,000-token threshold.
Introductory windows add a further wrinkle: Claude Sonnet 5's $2.00/$10.00 rate, as recorded by CloudZero in August 2026, is explicitly flagged by Spheron as an introductory price that runs only through 31 August 2026, with a distinct standard rate beginning the following day. This is exactly why a dated, sourced tracker matters more than a single headline price — comparison pages capture a snapshot, not a live feed, and vendors don't always publish a changelog when a rate quietly changes.
The Definitive August 2026 Price Comparison Table
The table below compiles the clearest current figures across budget, mid, and flagship tiers, each labelled with its source date. Note that Simplifai's April 2026 figures for Claude Sonnet 4 ($3.00/$15.00) and Claude Opus 4 ($15.00/$75.00) belong to an earlier model generation than the August 2026 Sonnet 5 and Opus 5 figures below, and should not be read as a price cut on the same product — they are different models entirely.
- •gpt-5.4-nano: $0.20 input / $1.25 output per million tokens (VibeEngines, July 2026)
- •Gemini 2.5 Flash-Lite: $0.10 input / $0.40 output per million tokens (VibeEngines, July 2026)
- •Claude Haiku 4.5: $1.00 input / $5.00 output per million tokens (CloudZero, August 2026)
- •Gemini 3.1 Pro (≤200K tokens): $2.00 input / $12.00 output per million tokens (Spheron, August 2026)
- •Claude Sonnet 5 (introductory, through 31 Aug 2026): $2.00 input / $10.00 output per million tokens (CloudZero, August 2026)
- •Claude Opus 5: $5.00 input / $25.00 output per million tokens — see the Anthropic Opus 5 pricing details (CloudZero, August 2026)
- •Claude Fable 5: $10.00 input / $50.00 output per million tokens (CloudZero, August 2026)
The Widening Spread Between Budget and Flagship Tiers
Lining these figures up side by side shows a spread rather than a single 'AI price', with output rates ranging from $0.40 per million tokens at the cheapest budget tier to $50.00 per million tokens at the most expensive flagship tier recorded in August 2026.
There is also a generational story hiding in the dataset: Simplifai's April 2026 figures for Claude Sonnet 4 ($3.00/$15.00) and Opus 4 ($15.00/$75.00) sit noticeably above CloudZero's August 2026 figures for Sonnet 5 ($2.00/$10.00 introductory) and Opus 5 ($5.00/$25.00). Read carefully, that's a newer generation launching at a lower headline price than its predecessor, rather than a mid-cycle discount on an unchanged product — an important distinction when a procurement team is deciding whether a 'price cut' is real.
Beyond the Sticker Price: What UK Buyers Must Weigh
List prices in this tracker are all published in US dollars. For UK procurement, that means every quoted rate needs converting at the invoice-date exchange rate before it goes into a budget model, since sterling movement against the dollar changes the effective cost per million tokens independently of anything the vendor does. VAT treatment is a separate line item procurement teams need to confirm directly with each vendor and their own accounting team rather than assume from a headline price.
Data residency and compliance sit alongside the rate card. UK GDPR obligations, sector-specific assurance processes, and public-sector procurement rules can rule out a cheaper model outright if it cannot demonstrate appropriate data handling — in which case the comparison exercise is really about which compliant vendors remain, and only then about which of those is cheapest per million tokens.
Rate limits and context-window ceilings also affect real-world cost: a model that is cheap per token but forces more round trips, more retries, or smaller context windows can end up costing more in engineering time and latency than a nominally pricier model that handles a task in a single call.
View the data behind this chart
| gpt-5.4-nano | Gemini 2.5 Flash-Lite | Claude Haiku 4.5 | Claude Opus 5 | Claude Fable 5 | |
|---|---|---|---|---|---|
| Input $/1M tokens | $/1M tok…0.2 | $/1M tok…0.1 | $/1M tok…1 | $/1M tok…5 | $/1M tok…10 |
| Output $/1M tokens | $/1M tok…1.25 | $/1M tok…0.4 | $/1M tok…5 | $/1M tok…25 | $/1M tok…50 |
Choosing the Right Model: A Strategic Framework for UK Teams
Task complexity should set the tier before price does. Three illustrative worked examples, using the verified per-token rates above and round assumed monthly volumes to show how the maths scales:
A support chatbot handling routine queries, illustratively assumed at 5 million input tokens and 1 million output tokens a month, would cost roughly $0.90 a month on Gemini 2.5 Flash-Lite ($0.10/$0.40) versus roughly $2.25 a month on gpt-5.4-nano ($0.20/$1.25) — both budget-tier options suited to low-complexity, high-volume work.
A content team generating longer copy, illustratively assumed at 2 million input tokens and 2 million output tokens a month, would cost roughly $28 a month on Gemini 3.1 Pro at its ≤200K-token rate ($2.00/$12.00), versus roughly $24 a month on Claude Sonnet 5's introductory rate ($2.00/$10.00) — noting that Sonnet 5's rate changes after 31 August 2026.
A small development team running a code assistant, illustratively assumed at 1 million input tokens and 0.5 million output tokens a month, would cost roughly $17.50 a month on Claude Opus 5 ($5.00/$25.00) versus roughly $35 a month on the top-tier Claude Fable 5 ($10.00/$50.00) — the flagship tier reserved for the hardest reasoning tasks.
For teams weighing self-hosted open-source models instead of any API tier, it's worth reading how the infrastructure and expertise costs stack up before committing — compare self-hosting LLMs with cloud GPU costs and review UK GPU cloud rental prices as a sanity check against the API rates above.
Optimising Spend: Context Windows, Introductory Rates and the Price War Ahead
The long-context tax is one of the clearest optimisation levers in this dataset: Gemini 2.5 Pro's input rate rises from $1.25 to $2.50 per million tokens once a prompt crosses 200,000 tokens, and Gemini 3.1 Pro's input and output rates both roughly double above the same threshold. Keeping prompts under that line, where the workload allows it, is a direct and measurable saving.
Time-limited introductory pricing is the other lever to watch. Claude Sonnet 5's $2.00/$10.00 rate is explicitly an introductory price through 31 August 2026, with a standard rate following on 1 September 2026 — budgets built around the introductory figure need a review point at that date rather than an assumption that the rate is permanent.
Taken together, the pattern across this dataset — cheaper budget-tier entries from Google and OpenAI, a widening gap to flagship reasoning tiers, and generation-on-generation launches at lower headline prices than their predecessors — points to continued downward pressure at the low end through late 2026, even as flagship reasoning tiers hold a premium. Buyers should treat any single quarter's table as a snapshot, not a forecast, and revisit it before renewing annual commitments.
Methodology
This tracker compiles per-token pricing figures published on independent LLM pricing comparison and tracker sites between April and August 2026, rather than scraping vendor billing consoles directly. Sources used are Simplifai (April 2026), SaaSTweaks (May 2026), VibeEngines (July 2026), and CloudZero and Spheron (both August 2026).
Each figure retains the scope stated by its original source — the model name, the input or output rate, the context-length band it applies to, and whether it was described as an introductory or standard rate. Where two sources reported different generations of the same model family (for example Claude Sonnet 4 versus Sonnet 5), those figures were kept separate rather than merged, since they describe different products at different points in time.
Figures were cross-checked across at least two independently dated snapshots where available, and every number in this article is reproduced exactly as published by its source, with no averaging, extrapolation, or currency conversion applied beyond what the source itself stated.
Sources
Every figure in this article traces to the sources below.
- •Simplifai — 2026 LLM API pricing comparison (GPT-4o mini, Claude Sonnet 4, Claude Opus 4)
- •SaaSTweaks — Gemini 2.5 Pro tiered pricing below/above 200K tokens
- •CloudZero — August 2026 Anthropic pricing snapshot (Haiku 4.5, Sonnet 5, Opus 5, Fable 5)
- •Spheron — August 2026 Gemini 3.1 Pro pricing and Claude Sonnet 5 introductory-rate timeline
- •VibeEngines — July 2026 LLM API pricing handbook (Gemini 2.5 Flash-Lite, gpt-5.4-nano, Claude Haiku 4.5)
The 10 verified data points behind this study are free to download and reuse with attribution (CC BY 4.0).
Cite as: Servnet Research, “LLM API Price Tracker 2026: Cost Per Million Tokens”, servnetuk.com, 2026.