Anthropic's Friday launch of Opus 5 at half the token price of its Fable 5 sibling isn't just a vendor discount — it's a signal that the token-cost war is now shaping how UK enterprises should model inference spend across cloud APIs and on-premise AI inference for the rest of 2026.
View the data behind this chart
| Fable 5 | Opus 5 (max) | GPT-5.6 Sol (max) | Kimi K3 | |
|---|---|---|---|---|
| Cost per task | $2.75 | $2.03 | $1.04 | $0.95 |
What Anthropic actually changed
Opus 5 is priced at $5 per million input tokens and $25 per million output tokens, exactly half of Fable 5's $10/$50 per million rate, per The Register's coverage of Friday's launch. Anthropic claims Opus 5 "comes close to the frontier intelligence of Claude Fable 5 at half the price." For context, OpenAI's GPT-5.6 Sol sits at $5/$30 per million tokens, meaning Opus 5 now undercuts a direct rival on output pricing while matching it on input.
This matters for procurement teams building 2026 budgets around per-model rate cards. But rate cards alone rarely tell the whole cost story, which is where Anthropic's own comparison metric — and Artificial Analysis's task-based benchmark — becomes more useful for UK buyers modelling real workloads.
Token price isn't the number that matters most
Per Artificial Analysis, the weighted average cost per Intelligence Index task is $2.75 for Fable 5, $2.03 for Opus 5 (max), $1.04 for GPT-5.6 Sol (max), and $0.95 for Kimi K3. Opus 5 leads the Intelligence Index at 61, one point ahead of Fable 5. For buyers, this is the more honest comparison: a model that charges less per token but needs more tokens to finish a task can still cost more in production than a pricier model that reasons in fewer steps.
This confirms what The Register argued in its own separate analysis of AI cost calculation — that task completion rates, not headline token prices, should drive procurement decisions. UK enterprises running high-volume agentic workloads should benchmark actual task cost before locking into annual API contracts.
The system prompt diet changes effective context economics
Anthropic has removed roughly 80 percent of the Claude Code system prompt for its latest models, according to Claude Code engineer Thariq Shihipar, freeing up more context window for customer prompts, skills, and CLAUDE.md files. Shihipar has advised customers to simplify their own prompt and skills configurations "just like we did," and pointed to a claude doctor command that can automatically optimise some of this work.
For UK teams running large context-window workloads — legal document review, codebase-wide agentic tasks, long customer transcripts — this effectively increases usable context per pound spent, without Anthropic changing the advertised token price. It's a genuine TCO lever buyers should factor into 2026 capacity planning, particularly when optimising purchases for AI inference workloads rather than training.

Verbosity is a hidden cost line
Anthropic itself warns that "Claude Opus 5's default user-facing responses run longer than prior Opus models'," and advises customers who want tighter output to "prompt for it explicitly." Because output tokens cost five times more than input tokens under Opus 5's pricing ($25 vs $5 per million), unmanaged verbosity could quietly erode the headline 50 percent saving over Fable 5.
Finance and platform teams should treat prompt engineering for response length as a cost-control task, not just a UX preference, when rolling Opus 5 into production pipelines.
Safety posture affects which UK sectors can actually deploy it
Opus 5 scores close to Anthropic's restricted Mythos model on finding vulnerabilities in open source code — 80 percent on the OSS-Fuzz benchmark versus 79.4 percent — but is markedly weaker at exploiting those findings, succeeding in 4 of 14 attempts compared with 13 of 14 for Mythos. Anthropic says Opus 5's cyber classifiers are "proportionally less restrictive" than Fable 5's, permitting vulnerability discovery in source code while still blocking binary-based scanning, penetration testing, and exploit generation.
Anthropic also describes Opus 5 as its "most aligned model to date" and is rolling out an automated fallback function so overly cautious refusals route to a less capable model rather than being blocked outright. For regulated UK buyers in finance, health and public sector, the more compelling commercial detail may be the lack of a data retention requirement on Opus 5 — a governance point that can simplify vendor risk assessments compared with retention policies attached to other Anthropic model tiers.
View the data behind this chart
| Opus 5 | Fable 5 | GPT-5.6 Sol | |
|---|---|---|---|
| Input price /MTok | $5 | $10 | $5 |
| Output price /MTok | $25 | $50 | $30 |
| Cost per Index task | $2.03 | $2.75 | $1.04 |
| Intelligence Index score | 61 | 60 | Not disclosed |
Cloud API versus on-premise: the real 2026 TCO decision
Opus 5's pricing pressure sits inside a wider market squeeze from open-weight rivals, and NVIDIA's own total-cost-of-ownership research on fine-tuning large language models points to dramatically lower cost-per-million-tokens when workloads are optimised on newer accelerator platforms rather than run on ageing hardware. That's the crux of the buy-versus-build question UK infrastructure leads face this year.
API pricing from Anthropic and OpenAI is falling, but so is the entry cost of self-hosting comparable open-weight models. Teams with predictable, high-volume inference workloads should calculate your cloud vs. on-premise TCO before renewing enterprise API contracts, and check GPU server requirements for LLM inference against current accelerator pricing. For a broader market baseline, the latest AI server cost index for 2026 and a dedicated look at self-hosting LLM vs. cloud GPU costs in the UK are useful reference points for building the business case either way.
- 01The Register — Anthropic debuts Opus 5 at half the price of its Fable sibling · 25 July 2026
- 02The Register — The price is wrong: AI cost calculation has to consider task completion rates · 13 July 2026
- 03TechRadar — China's answer to Claude's Fable 5 tops HTML web design contest · 10 July 2026
- 04NVIDIA — TCO comparison: fine-tuning large language models on accelerator platforms · 1 June 2026
