UK’s trusted IT infrastructure partner since 2003
Servnet
FinanceToolsConfiguratorGet in Touch
AI Infrastructure

LLM API Price Tracker 2026: Cost Per Million Tokens

Servnet Editorial · IT infrastructure analysis7 min read
Share

As of 23 August 2026, the cheapest published output price among major LLM APIs sits at $0.40 per million output tokens for Google's Gemini 2.5 Flash-Lite, while Anthropic's flagship Claude Fable 5 is listed at $50.00 per million output tokens — two entries in the same live pricing landscape. Between those extremes sit Anthropic's Claude Haiku 4.5, Sonnet 5 and Opus 5, Google's Gemini 3.1 Pro, and OpenAI's gpt-5.4-nano, each dated to the comparison page that published it, because vendors increasingly reprice, add context-length tiers, or run time-limited introductory rates without a formal changelog. This tracker compiles every verifiable per-token figure published between April and August 2026, dates each one to its source, and works through what the numbers mean for UK IT buyers weighing build-vs-buy decisions.

August 2026 Price Snapshot by Model
Input $/1M…Output $/1M…Source dategpt-5.4-nano$0.20$1.25Jul 2026Gemini 2.5 Flash-Lite$0.10$0.40Jul 2026Claude Haiku 4.5$1.00$5.00Aug 2026Gemini 3.1 Pro (≤200K)$2.00$12.00Aug 2026Claude Sonnet 5…$2.00$10.00Aug 2026Claude Opus 5$5.00$25.00Aug 2026Claude Fable 5$10.00$50.00Aug 2026
View the data behind this chart
August 2026 Price Snapshot by Model
Input $/1M…Output $/1M…Source date
gpt-5.4-nano$0.20$1.25Jul 2026
Gemini 2.5 Flash-Lite$0.10$0.40Jul 2026
Claude Haiku 4.5$1.00$5.00Aug 2026
Gemini 3.1 Pro (≤200K)$2.00$12.00Aug 2026
Claude Sonnet 5…$2.00$10.00Aug 2026
Claude Opus 5$5.00$25.00Aug 2026
Claude Fable 5$10.00$50.00Aug 2026

August 2026 Snapshot: What Actually Changed

CloudZero's August 2026 comparison lists Anthropic's current line-up as Claude Haiku 4.5 at $1.00 per million input tokens and $5.00 per million output tokens, Claude Sonnet 5 at $2.00/$10.00, Claude Opus 5 at $5.00/$25.00, and the top-tier Claude Fable 5 at $10.00/$50.00 per million tokens.

Separately, Spheron's August 2026 tracker put Google's Gemini 3.1 Pro at $2.00 input and $12.00 output per million tokens for prompts at or below 200,000 tokens, rising to $4.00/$18.00 above that threshold — and flagged that Claude Sonnet 5's pricing is an introductory rate running through 31 August 2026, with a separate standard rate scheduled to begin 1 September 2026.

At the budget end, VibeEngines' July 2026 handbook recorded Gemini 2.5 Flash-Lite at $0.10 input / $0.40 output per million tokens — the cheapest entry-tier figure in this dataset — alongside OpenAI's gpt-5.4-nano at $0.20/$1.25 and Anthropic's Claude Haiku 4.5 at $1.00/$5.00, the same Haiku figure independently confirmed a month later by CloudZero's August snapshot.

Illustration: LLM API Price Tracker 2026: Cost Per Million Tokens

How Token Pricing Actually Works: Tiers, Thresholds and Quiet Repricing

Every provider in this dataset charges output tokens at a materially higher rate than input tokens — a pattern worth checking before estimating cost from input volume alone. Readers new to the mechanics may find it useful to understand LLM tokens and context windows before comparing rate cards, since 'a million tokens' is not the same unit of work across every model or task.

Context length is now part of the price, not just a spec. SaaSTweaks' May 2026 comparison shows Google's Gemini 2.5 Pro billed at $1.25 per million input tokens for prompts under 200,000 tokens, rising to $2.50 above that line. Spheron's August 2026 figures show the same pattern for Gemini 3.1 Pro, where both input and output rates roughly double once a prompt crosses the 200,000-token threshold.

Introductory windows add a further wrinkle: Claude Sonnet 5's $2.00/$10.00 rate, as recorded by CloudZero in August 2026, is explicitly flagged by Spheron as an introductory price that runs only through 31 August 2026, with a distinct standard rate beginning the following day. This is exactly why a dated, sourced tracker matters more than a single headline price — comparison pages capture a snapshot, not a live feed, and vendors don't always publish a changelog when a rate quietly changes.

The Definitive August 2026 Price Comparison Table

The table below compiles the clearest current figures across budget, mid, and flagship tiers, each labelled with its source date. Note that Simplifai's April 2026 figures for Claude Sonnet 4 ($3.00/$15.00) and Claude Opus 4 ($15.00/$75.00) belong to an earlier model generation than the August 2026 Sonnet 5 and Opus 5 figures below, and should not be read as a price cut on the same product — they are different models entirely.

  • gpt-5.4-nano: $0.20 input / $1.25 output per million tokens (VibeEngines, July 2026)
  • Gemini 2.5 Flash-Lite: $0.10 input / $0.40 output per million tokens (VibeEngines, July 2026)
  • Claude Haiku 4.5: $1.00 input / $5.00 output per million tokens (CloudZero, August 2026)
  • Gemini 3.1 Pro (≤200K tokens): $2.00 input / $12.00 output per million tokens (Spheron, August 2026)
  • Claude Sonnet 5 (introductory, through 31 Aug 2026): $2.00 input / $10.00 output per million tokens (CloudZero, August 2026)
  • Claude Opus 5: $5.00 input / $25.00 output per million tokens — see the Anthropic Opus 5 pricing details (CloudZero, August 2026)
  • Claude Fable 5: $10.00 input / $50.00 output per million tokens (CloudZero, August 2026)

The Widening Spread Between Budget and Flagship Tiers

Lining these figures up side by side shows a spread rather than a single 'AI price', with output rates ranging from $0.40 per million tokens at the cheapest budget tier to $50.00 per million tokens at the most expensive flagship tier recorded in August 2026.

There is also a generational story hiding in the dataset: Simplifai's April 2026 figures for Claude Sonnet 4 ($3.00/$15.00) and Opus 4 ($15.00/$75.00) sit noticeably above CloudZero's August 2026 figures for Sonnet 5 ($2.00/$10.00 introductory) and Opus 5 ($5.00/$25.00). Read carefully, that's a newer generation launching at a lower headline price than its predecessor, rather than a mid-cycle discount on an unchanged product — an important distinction when a procurement team is deciding whether a 'price cut' is real.

Beyond the Sticker Price: What UK Buyers Must Weigh

List prices in this tracker are all published in US dollars. For UK procurement, that means every quoted rate needs converting at the invoice-date exchange rate before it goes into a budget model, since sterling movement against the dollar changes the effective cost per million tokens independently of anything the vendor does. VAT treatment is a separate line item procurement teams need to confirm directly with each vendor and their own accounting team rather than assume from a headline price.

Data residency and compliance sit alongside the rate card. UK GDPR obligations, sector-specific assurance processes, and public-sector procurement rules can rule out a cheaper model outright if it cannot demonstrate appropriate data handling — in which case the comparison exercise is really about which compliant vendors remain, and only then about which of those is cheapest per million tokens.

Rate limits and context-window ceilings also affect real-world cost: a model that is cheap per token but forces more round trips, more retries, or smaller context windows can end up costing more in engineering time and latency than a nominally pricier model that handles a task in a single call.

Input vs Output Price by Model, August 2026
$/1M tok…50$/1M tok…38$/1M tok…25$/1M tok…13$/1M tok…0$/1M tok…0.2$/1M tok…1.25gpt-5.4-nano$/1M tok…0.1$/1M tok…0.4Gemini 2.5Flash-Lite$/1M tok…1$/1M tok…5Claude Haiku 4.5$/1M tok…5$/1M tok…25Claude Opus 5$/1M tok…10$/1M tok…50Claude Fable 5Input $/1M tokensOutput $/1M tokens
View the data behind this chart
Input vs Output Price by Model, August 2026
gpt-5.4-nanoGemini 2.5 Flash-LiteClaude Haiku 4.5Claude Opus 5Claude Fable 5
Input $/1M tokens$/1M tok…0.2$/1M tok…0.1$/1M tok…1$/1M tok…5$/1M tok…10
Output $/1M tokens$/1M tok…1.25$/1M tok…0.4$/1M tok…5$/1M tok…25$/1M tok…50

Choosing the Right Model: A Strategic Framework for UK Teams

Task complexity should set the tier before price does. Three illustrative worked examples, using the verified per-token rates above and round assumed monthly volumes to show how the maths scales:

A support chatbot handling routine queries, illustratively assumed at 5 million input tokens and 1 million output tokens a month, would cost roughly $0.90 a month on Gemini 2.5 Flash-Lite ($0.10/$0.40) versus roughly $2.25 a month on gpt-5.4-nano ($0.20/$1.25) — both budget-tier options suited to low-complexity, high-volume work.

A content team generating longer copy, illustratively assumed at 2 million input tokens and 2 million output tokens a month, would cost roughly $28 a month on Gemini 3.1 Pro at its ≤200K-token rate ($2.00/$12.00), versus roughly $24 a month on Claude Sonnet 5's introductory rate ($2.00/$10.00) — noting that Sonnet 5's rate changes after 31 August 2026.

A small development team running a code assistant, illustratively assumed at 1 million input tokens and 0.5 million output tokens a month, would cost roughly $17.50 a month on Claude Opus 5 ($5.00/$25.00) versus roughly $35 a month on the top-tier Claude Fable 5 ($10.00/$50.00) — the flagship tier reserved for the hardest reasoning tasks.

For teams weighing self-hosted open-source models instead of any API tier, it's worth reading how the infrastructure and expertise costs stack up before committing — compare self-hosting LLMs with cloud GPU costs and review UK GPU cloud rental prices as a sanity check against the API rates above.

Optimising Spend: Context Windows, Introductory Rates and the Price War Ahead

The long-context tax is one of the clearest optimisation levers in this dataset: Gemini 2.5 Pro's input rate rises from $1.25 to $2.50 per million tokens once a prompt crosses 200,000 tokens, and Gemini 3.1 Pro's input and output rates both roughly double above the same threshold. Keeping prompts under that line, where the workload allows it, is a direct and measurable saving.

Time-limited introductory pricing is the other lever to watch. Claude Sonnet 5's $2.00/$10.00 rate is explicitly an introductory price through 31 August 2026, with a standard rate following on 1 September 2026 — budgets built around the introductory figure need a review point at that date rather than an assumption that the rate is permanent.

Taken together, the pattern across this dataset — cheaper budget-tier entries from Google and OpenAI, a widening gap to flagship reasoning tiers, and generation-on-generation launches at lower headline prices than their predecessors — points to continued downward pressure at the low end through late 2026, even as flagship reasoning tiers hold a premium. Buyers should treat any single quarter's table as a snapshot, not a forecast, and revisit it before renewing annual commitments.

Methodology

This tracker compiles per-token pricing figures published on independent LLM pricing comparison and tracker sites between April and August 2026, rather than scraping vendor billing consoles directly. Sources used are Simplifai (April 2026), SaaSTweaks (May 2026), VibeEngines (July 2026), and CloudZero and Spheron (both August 2026).

Each figure retains the scope stated by its original source — the model name, the input or output rate, the context-length band it applies to, and whether it was described as an introductory or standard rate. Where two sources reported different generations of the same model family (for example Claude Sonnet 4 versus Sonnet 5), those figures were kept separate rather than merged, since they describe different products at different points in time.

Figures were cross-checked across at least two independently dated snapshots where available, and every number in this article is reproduced exactly as published by its source, with no averaging, extrapolation, or currency conversion applied beyond what the source itself stated.

Sources

Every figure in this article traces to the sources below.

  • Simplifai — 2026 LLM API pricing comparison (GPT-4o mini, Claude Sonnet 4, Claude Opus 4)
  • SaaSTweaks — Gemini 2.5 Pro tiered pricing below/above 200K tokens
  • CloudZero — August 2026 Anthropic pricing snapshot (Haiku 4.5, Sonnet 5, Opus 5, Fable 5)
  • Spheron — August 2026 Gemini 3.1 Pro pricing and Claude Sonnet 5 introductory-rate timeline
  • VibeEngines — July 2026 LLM API pricing handbook (Gemini 2.5 Flash-Lite, gpt-5.4-nano, Claude Haiku 4.5)
Open data

The 10 verified data points behind this study are free to download and reuse with attribution (CC BY 4.0).

Cite as: Servnet Research, “LLM API Price Tracker 2026: Cost Per Million Tokens”, servnetuk.com, 2026.

Share
Key takeaways
  • Cheapest budget-tier entries differ by provider: Gemini 2.5 Flash-Lite at $0.10/$0.40 vs gpt-5.4-nano at $0.20/$1.25 per million tokens (both July 2026).
  • Flagship-tier spread is wide: Claude Opus 5 at $5.00/$25.00 vs Claude Fable 5 at $10.00/$50.00 per million tokens (both August 2026, CloudZero).
  • Context length carries a direct price tax: Gemini 3.1 Pro's input rate doubles from $2.00 to $4.00 per million tokens above 200,000 tokens (Spheron, August 2026).
  • Claude Sonnet 5's $2.00/$10.00 rate is an introductory price running only through 31 August 2026, with a standard rate beginning 1 September 2026 (Spheron).
  • UK buyers must add FX conversion, VAT treatment, and data-residency/compliance checks on top of every US-dollar list price before comparing vendors.
  • Generation-on-generation launches (Sonnet 4/Opus 4 in April vs Sonnet 5/Opus 5 in August) show lower headline prices at launch — treat this as a new product, not a discount.
Frequently asked

FAQs — LLM API Price Tracker 2026

What is the cheapest LLM API per million tokens as of August 2026?

The cheapest entry-tier figure in this dataset is Google's Gemini 2.5 Flash-Lite at $0.10 per million input tokens and $0.40 per million output tokens, recorded by VibeEngines' July 2026 tracker. OpenAI's gpt-5.4-nano and Anthropic's Claude Haiku 4.5 sit slightly above it in the same budget category.

Is Claude Sonnet 5's pricing changing soon?

Yes. CloudZero's August 2026 snapshot lists Claude Sonnet 5 at $2.00 input / $10.00 output per million tokens, but Spheron's August 2026 tracker flags this as an introductory rate that runs only through 31 August 2026, with a separate standard rate beginning 1 September 2026.

Why do Gemini prices change above 200,000 tokens?

Google prices Gemini 2.5 Pro and Gemini 3.1 Pro on a two-tier basis: a lower rate for prompts at or under 200,000 tokens and a higher rate above it. Gemini 2.5 Pro's input rate rises from $1.25 to $2.50 per million tokens; Gemini 3.1 Pro's input and output rates both roughly double above the same threshold.

Do UK businesses need to worry about VAT on LLM API bills?

Published API prices are quoted in US dollars and typically exclude UK VAT considerations entirely. UK buyers need to confirm VAT treatment directly with each vendor and their finance team, and should factor exchange-rate movement into any budget built around dollar-denominated per-token rates.

Should a UK business self-host an open-source model instead of using an API?

It depends on volume, expertise, and compliance needs. API pricing avoids infrastructure and maintenance overhead but scales with usage, while self-hosting shifts cost to hardware and specialist staff time. Comparing the two properly requires looking beyond the per-token rate to total infrastructure and operational cost.

Why do LLM providers reprice without an announcement?

Several tracked comparisons in this dataset show prices changing between one dated snapshot and the next — for example Claude Sonnet 5's introductory-to-standard transition — without a dedicated vendor changelog entry. This is why a quarterly, dated tracker is more reliable than a single cached price page.

Related

Continue reading

More in Research

Got a question this study didn't answer?

One conversation with an engineer who's done this before. No sales script.

Talk to Servnet →

Talk to a UK specialist

Get expert advice or a no-obligation quote — servers, storage, networking, maintenance, finance and cloud. We reply the same working day.

or call 0800 987 4111