Majestic Labs has unveiled Prometheus, a server that ditches Nvidia GPUs and HBM in favour of Arm-based processors and up to 128TB of pooled LPDDR6 memory. For UK buyers wrestling with GPU scarcity and cost, it's a serious enough proposition to warrant scrutiny — but not yet a signed contract.
View the data behind this chart
| DGX B300 | HPE 8-GPU node | Prometheus max | |
|---|---|---|---|
| Fast memory (TB) | TB2.3 | TB8 | TB128 |
What Majestic Labs is actually proposing
The Tel Aviv-founded startup, built by former Google and Meta engineers, argues that bolting expensive GPUs to high-bandwidth memory has become a dead end for inference workloads because performance is now bound by memory access, not raw compute. Its answer is Prometheus, a server built around custom Ignite AI Processing Units that combine Arm cores with RISC-V vector and tensor engines rather than Nvidia silicon.
Each server can carry up to 12 of these AIUs, sharing a single coherent memory pool of between 8TB and 128TB of LPDDR6 — reached via custom memory aggregation chiplets over copper cabling up to a metre long, rather than memory soldered directly onto GPU packages. A 40U rack holds four such servers, drawing 120kW total and cooled with cold-plate liquid systems rather than air. UK buyers who compare air vs. liquid cooling for AI servers will recognise this as a meaningful facilities commitment, not a drop-in replacement for an existing air-cooled hall.
The memory-wall argument, and why it matters here
Majestic's central claim is stark: over 50 times more fast memory than a comparable Nvidia DGX B300 configuration, which offers 2.3TB of HBM3e plus up to 4TB of DDR5 system memory, alongside 1.7 times the interconnect bandwidth. The company also claims one Majestic rack can match the fast memory capacity of 25 Nvidia NVL72 Vera Rubin racks at a fraction of the power draw.
For context, mainstream accelerated servers from established vendors are still shipping with configurations such as eight Nvidia H200, B200, B300, or AMD MI355X accelerators, with system memory ceilings around 8TB on some builds — a fraction of Prometheus's claimed 128TB pool. Buyers who have already had to understand the HBM memory crunch driving up GPU-attached hardware costs will see immediately why a memory-pooled architecture is an attractive pitch, at least on paper.
Cost-per-token: the number that actually matters
Majestic says Prometheus could cost between 10 and 50 times less than a GPU system of equivalent performance once it ships next year, while consuming less electricity per rack. That's the figure UK procurement teams should interrogate hardest, because list-price comparisons rarely survive contact with real inference workloads, batching efficiency, and software maturity.
The server is pitched as OCP-compliant and designed to run PyTorch, vLLM and OpenAI's Triton without modification — a sensible move to lower the switching barrier for teams that have already optimise for AI inference workloads on existing frameworks. But software compatibility on paper is not the same as production-grade performance parity, particularly for latency-sensitive or high-throughput serving.

What's still unverified
Several engineering questions remain open. TechRadar Pro's analysis notes that if a 128TB configuration is built from widely available 2GB LPDDR6 dies, a single server would need roughly 64,000 of them, implying well over a hundred memory aggregation chiplets — a level of integration complexity that has yet to be demonstrated at scale or independently benchmarked.
Majestic says it has already secured orders from large enterprises, neoclouds and hyperscalers, and has raised $100 million in an A-round, which suggests some buyers are willing to bet early. But every performance and cost figure quoted so far is the startup's own projection ahead of shipped hardware or third-party testing. UK buyers should treat these as directional claims, not procurement-grade benchmarks, until independent validation appears.
How UK buyers should approach this decision
Nothing here changes near-term reality: Nvidia GPUs remain the default, and Arm-based server CPUs are already gaining traction elsewhere in the market as evidence that high-core-count Arm silicon is becoming credible for AI-adjacent workloads. But for organisations priced out of hyperscaler-grade GPU clusters, a memory-centric alternative shipping next year is worth tracking now, not after launch.
Practical next steps for procurement and infrastructure teams: use current-generation GPU pricing and availability to assess GPU requirements and costs against your actual inference workload profile before assuming any alternative architecture will be cheaper in practice; and where budget cycles allow, explore how you might finance new AI server deployments flexibly enough to pivot if a validated alternative to GPU-based inference emerges in 2026 and 2027. Teams new to this category should also learn more about inference servers before committing capital either way.
- 01TechRadar Pro — Startup swaps costly AI GPUs for Arm cores and up to 128TB of cheap LPDDR6 RAM instead of expensive HBM to smash through the memory wall · 1 August 2026
- 02The Next Platform — Arm comes full circle with homegrown AI-tuned server CPU · 25 March 2026
- 03NVIDIA — RTX PRO AI Factory reference architecture components · 1 January 2026
- 04HPE — AI servers · 1 January 2026
