UK’s trusted IT infrastructure partner since 2003
Servnet
FinanceToolsConfiguratorGet in Touch
AI / GPU

Anthropic Opus 5 Pricing 2026: UK Buyer Playbook

London · Servnet News Desk · IT infrastructure analysis4 min read
Share

Anthropic's Friday launch of Opus 5 at half the token price of its Fable 5 sibling isn't just a vendor discount — it's a signal that the token-cost war is now shaping how UK enterprises should model inference spend across cloud APIs and on-premise AI inference for the rest of 2026.

Weighted average cost per Intelligence Index task
$10$8$5$3$0$2.75Fable 5$2.03Opus 5 (max)$1.04GPT-5.6 Sol (max)$0.95Kimi K3Cost per task
View the data behind this chart
Weighted average cost per Intelligence Index task
Fable 5Opus 5 (max)GPT-5.6 Sol (max)Kimi K3
Cost per task$2.75$2.03$1.04$0.95

What Anthropic actually changed

Opus 5 is priced at $5 per million input tokens and $25 per million output tokens, exactly half of Fable 5's $10/$50 per million rate, per The Register's coverage of Friday's launch. Anthropic claims Opus 5 "comes close to the frontier intelligence of Claude Fable 5 at half the price." For context, OpenAI's GPT-5.6 Sol sits at $5/$30 per million tokens, meaning Opus 5 now undercuts a direct rival on output pricing while matching it on input.

This matters for procurement teams building 2026 budgets around per-model rate cards. But rate cards alone rarely tell the whole cost story, which is where Anthropic's own comparison metric — and Artificial Analysis's task-based benchmark — becomes more useful for UK buyers modelling real workloads.

Token price isn't the number that matters most

Per Artificial Analysis, the weighted average cost per Intelligence Index task is $2.75 for Fable 5, $2.03 for Opus 5 (max), $1.04 for GPT-5.6 Sol (max), and $0.95 for Kimi K3. Opus 5 leads the Intelligence Index at 61, one point ahead of Fable 5. For buyers, this is the more honest comparison: a model that charges less per token but needs more tokens to finish a task can still cost more in production than a pricier model that reasons in fewer steps.

This confirms what The Register argued in its own separate analysis of AI cost calculation — that task completion rates, not headline token prices, should drive procurement decisions. UK enterprises running high-volume agentic workloads should benchmark actual task cost before locking into annual API contracts.

The system prompt diet changes effective context economics

Anthropic has removed roughly 80 percent of the Claude Code system prompt for its latest models, according to Claude Code engineer Thariq Shihipar, freeing up more context window for customer prompts, skills, and CLAUDE.md files. Shihipar has advised customers to simplify their own prompt and skills configurations "just like we did," and pointed to a claude doctor command that can automatically optimise some of this work.

For UK teams running large context-window workloads — legal document review, codebase-wide agentic tasks, long customer transcripts — this effectively increases usable context per pound spent, without Anthropic changing the advertised token price. It's a genuine TCO lever buyers should factor into 2026 capacity planning, particularly when optimising purchases for AI inference workloads rather than training.

Illustration: Anthropic Opus 5 Pricing 2026: UK Buyer Playbook

Verbosity is a hidden cost line

Anthropic itself warns that "Claude Opus 5's default user-facing responses run longer than prior Opus models'," and advises customers who want tighter output to "prompt for it explicitly." Because output tokens cost five times more than input tokens under Opus 5's pricing ($25 vs $5 per million), unmanaged verbosity could quietly erode the headline 50 percent saving over Fable 5.

Finance and platform teams should treat prompt engineering for response length as a cost-control task, not just a UX preference, when rolling Opus 5 into production pipelines.

Safety posture affects which UK sectors can actually deploy it

Opus 5 scores close to Anthropic's restricted Mythos model on finding vulnerabilities in open source code — 80 percent on the OSS-Fuzz benchmark versus 79.4 percent — but is markedly weaker at exploiting those findings, succeeding in 4 of 14 attempts compared with 13 of 14 for Mythos. Anthropic says Opus 5's cyber classifiers are "proportionally less restrictive" than Fable 5's, permitting vulnerability discovery in source code while still blocking binary-based scanning, penetration testing, and exploit generation.

Anthropic also describes Opus 5 as its "most aligned model to date" and is rolling out an automated fallback function so overly cautious refusals route to a less capable model rather than being blocked outright. For regulated UK buyers in finance, health and public sector, the more compelling commercial detail may be the lack of a data retention requirement on Opus 5 — a governance point that can simplify vendor risk assessments compared with retention policies attached to other Anthropic model tiers.

Opus 5 vs Fable 5 vs GPT-5.6 Sol
Opus 5Fable 5GPT-5.6 SolInput price /MTok$5$10$5Output price /MTok$25$50$30Cost per Index task$2.03$2.75$1.04Intelligence Index score6160Not disclosed
View the data behind this chart
Opus 5 vs Fable 5 vs GPT-5.6 Sol
Opus 5Fable 5GPT-5.6 Sol
Input price /MTok$5$10$5
Output price /MTok$25$50$30
Cost per Index task$2.03$2.75$1.04
Intelligence Index score6160Not disclosed

Cloud API versus on-premise: the real 2026 TCO decision

Opus 5's pricing pressure sits inside a wider market squeeze from open-weight rivals, and NVIDIA's own total-cost-of-ownership research on fine-tuning large language models points to dramatically lower cost-per-million-tokens when workloads are optimised on newer accelerator platforms rather than run on ageing hardware. That's the crux of the buy-versus-build question UK infrastructure leads face this year.

API pricing from Anthropic and OpenAI is falling, but so is the entry cost of self-hosting comparable open-weight models. Teams with predictable, high-volume inference workloads should calculate your cloud vs. on-premise TCO before renewing enterprise API contracts, and check GPU server requirements for LLM inference against current accelerator pricing. For a broader market baseline, the latest AI server cost index for 2026 and a dedicated look at self-hosting LLM vs. cloud GPU costs in the UK are useful reference points for building the business case either way.

Share
Key takeaways
  • Opus 5 costs half of Fable 5 per token ($5/$25 vs $10/$50 per million), but Artificial Analysis puts real task cost at $2.03 vs $2.75 — the smarter TCO metric for procurement.
  • An 80 percent cut to Claude Code's system prompt frees up context window, effectively lowering cost per usable token for large-context UK workloads.
  • Opus 5's default verbosity can erode savings since output tokens cost 5x input tokens; explicit length prompting is now a cost-control task, not just UX polish.
  • No data retention requirement and less restrictive cyber classifiers make Opus 5 easier to justify for regulated UK sectors than prior Anthropic tiers.
Frequently asked

FAQs — Anthropic Opus 5 Pricing 2026

Is Opus 5 actually cheaper than Fable 5 for UK enterprise workloads?

On token price, yes — Opus 5 is exactly half of Fable 5 at $5/$25 per million tokens versus $10/$50. On real task cost, Artificial Analysis puts Opus 5 (max) at $2.03 per Intelligence Index task versus $2.75 for Fable 5, still a meaningful saving.

How does Opus 5 compare with GPT-5.6 Sol on cost?

GPT-5.6 Sol prices at $5/$30 per million tokens and costs $1.04 per weighted Intelligence Index task, undercutting both Opus 5 and Fable 5 on task-based cost, per Artificial Analysis figures cited by The Register.

Does Opus 5's shorter system prompt actually save money?

Anthropic has cut roughly 80 percent of the Claude Code system prompt, freeing context window for customer prompts and skills, which can reduce the volume of tokens needed for equivalent tasks — though Anthropic advises customers to simplify their own prompts to see the full benefit.

Should UK enterprises consider on-premise inference instead of Anthropic's API?

It depends on workload volume and predictability. Buyers with steady, high-throughput inference needs should compare API pricing against self-hosted accelerator costs using tools like our AI GPU calculator before committing to either path for 2026.

Related

Continue reading

More in AI / GPU

Turning this into a buying decision?

One conversation with an engineer who's specced this before. No sales script.

Talk to Servnet →

Talk to a UK specialist

Get expert advice or a no-obligation quote — servers, storage, networking, maintenance, finance and cloud. We reply the same working day.

or call 0800 987 4111