Apple refreshed the Mac Studio on 25 August 2026 with M5 Max and M5 Ultra chips, and the number local-LLM buyers care about most is back: the M5 Ultra is configurable to 512GB of unified memory, with 1.2TB/s of memory bandwidth. That ends a lean spell in which, according to MacRumors, memory shortages had cut the M3 Ultra Mac Studio to a single 96GB configuration. But memory capacity only decides which models fit. How fast they run, and whether one box can serve a team, are separate questions, and the published tokens-per-second data is messier than most comparisons admit.
View the data behind this chart
| M2 Ultra alone | M2 Ultra + DGX Spark (RPC) | |
|---|---|---|
| Prompt tok/s | tok/s28.2 | tok/s29.5 |
| Generation tok/s | tok/s11.1 | tok/s5.9 |
Mac Studio for LLMs: The 2026 Verdict
The honest read on Mac Studio is narrower than the internet debate suggests. Its strength is capacity: one pool of unified memory big enough to hold a large quantized model without splitting it across several GPU cards. With the M5 Ultra now configurable to 256GB or 512GB, that argument is stronger than it has been all year. What the published benchmarks don't yet show is how fast the new chips run 70B-class models. The M5 models only reach customers from 22 September, and none of the sources we checked includes an M5 Ultra 70B result.
The older data sets expectations. On a 70B Q4_K_M generation test from May 2024, an Nvidia H100 PCIe 80GB produced 25.01 tok/s against 12.13 tok/s for an M2 Ultra with a 76-core GPU. That is roughly twice the generation speed on the same model class, so on raw speed Apple has no lead to defend. A March 2026 test pairing a 128GB M2 Ultra Mac Studio with an Nvidia DGX Spark over 10 Gbps Ethernet adds a second lesson: splitting a model across the two machines sped up prompt processing but slowed generation. The Mac's case rests on fitting big models privately on one desktop machine, not on out-running server GPUs.

Apple Silicon vs Nvidia: The Architectural Trade-off
Apple's pitch is unified memory capacity: one pool of RAM shared by CPU and GPU means a big quantized model can simply fit. On Apple's current UK specs page, the M5 Max Mac Studio starts at 36GB and is configurable to 48GB, 64GB or 128GB; the M5 Ultra starts at 96GB and is configurable to 256GB or 512GB, and Apple says the 512GB configuration arrives in late October. If you want to grasp why that matters before a model generates a single token, it's worth reading up on how much VRAM an LLM truly needs and taking the time to understand the difference between VRAM and RAM in the first place.
Capacity decides what fits; memory bandwidth largely decides generation speed, because every token a dense model generates means reading its weights. Apple lists 460GB/s for the 32-core-GPU M5 Max, 614GB/s for the 40-core-GPU M5 Max and 1.2TB/s for the M5 Ultra, which Apple describes as 50 percent higher than before; Apple listed the M3 Ultra at 819GB/s. A back-of-envelope ceiling makes this concrete: divide bandwidth by model size. The 44.2GB Qwen2.5-72B Q4_K_M file gives a theoretical ceiling of about 18 tok/s on the 800GB/s M2 Ultra, which measured 11.1 tok/s in practice, and about 27 tok/s on the 1.2TB/s M5 Ultra. Treat that as an upper bound for single-stream generation, not a benchmark: real runs land below it.
Tokens-per-Second Deep Dive: What the Benchmarks Actually Show
The most detailed Mac-plus-Nvidia data comes from a March 2026 GitHub benchmark of llama.cpp's RPC backend. It ran on a Mac Studio with an M2 Ultra and 128GB of unified memory, linked point-to-point over 10 Gbps Ethernet to a DGX Spark with Nvidia's GB10 Blackwell chip. On Qwen2.5-72B-Instruct at Q4_K_M, the Mac alone processed prompts at 28.2 tok/s and generated at 11.1 tok/s. Splitting the model (30.7GB on Metal, 13.8GB on CUDA) nudged prompt throughput up to 29.5 tok/s but cut generation to 5.9 tok/s.
The same repository's small-model test produced the eye-catching numbers, and they are easy to misread. On Qwen2.5-7B-Instruct Q4_K_M, a 4.4GB model rather than a 70B one, the Mac alone recorded 76.1 tok/s prompt and 91.8 tok/s generation, while the RPC split lifted prompt throughput to 317.7 tok/s and dropped generation to 52.7 tok/s. The pattern holds at both sizes: over a network link, split inference trades generation speed for prompt speed. The repository's own conclusion is that RPC's value is capacity, meaning models too big for either machine alone, rather than speed.
llama.cpp's long-running Apple Silicon results table adds a chip-to-chip view on the much smaller Llama 2 7B at Q4_0. The 60-core M3 Ultra recorded 1,073.09 tok/s prompt processing and 88.40 tok/s generation, the 40-core M4 Max 885.68 and 83.06, and the 40-core M5 Max, a chip option Apple also sells in the new Mac Studio, 3,219.99 and 119.92. On a model this small, compute matters as well as bandwidth, which is why the M5 Max out-generates the higher-bandwidth M3 Ultra here. Read these as relative chip capability only; a 7B result says little about absolute 70B speeds.
Choosing Your Mac Studio: Configurations and Value for UK Buyers
Apple's launch announcement and MacRumors' Mac Studio roundup put US starting prices at $2,499 for the M5 Max model and $5,499 for the M5 Ultra, with pre-orders open and availability from 22 September. We haven't verified UK sterling pricing for this article, so check Apple's UK store directly before budgeting: local pricing, VAT treatment and configuration availability move independently of the US figures.
For 70B-class models the memory tier matters more than the chip badge. The Qwen2.5-72B Q4_K_M file in the benchmark above is 44.2GB before any context is loaded, so a 64GB M5 Max leaves limited headroom for macOS, long prompts and the KV cache. The 128GB M5 Max and the M5 Ultra, from 96GB, are the realistic starting points, and the M5 Ultra's 1.2TB/s bandwidth should favour it for generation speed on the same model. The 256GB and 512GB M5 Ultra tiers are for models that won't fit smaller machines at all. That 512GB ceiling is the one Apple used in March 2025 to say an M3 Ultra Mac Studio could run LLMs with over 600 billion parameters entirely in memory. Availability is the practical caveat: the 512GB configuration isn't due until late October, and memory shortages have already removed top memory tiers once this year, so confirm lead times before you plan around them.
Worked Example: Getting a 70B Model Running, and the Software Stack That Matters
For a 70B-class model at Q4_K_M quantization on a 96GB or larger Mac Studio, the practical steps are: pick a runtime, pull a pre-quantized GGUF or MLX-format checkpoint, load it, and leave enough unified memory headroom above the model weights for context and system overhead before you start pushing long prompts through it.
Which runtime you pick genuinely changes the outcome, but the best published runtime comparison we found isn't a 70B test. Famstack's March 2026 write-up ran Qwen3 30B-A3B, a mixture-of-experts model, through an eight-turn ops-agent workload on a MacBook Pro with an M1 Max and 64GB, measuring effective tok/s, which counts prompt processing as well as generation. With GGUF weights on the llama.cpp engine, LM Studio managed 41.7, llama.cpp built from source 41.4 and Ollama 26.0, which is 37% slower than LM Studio on the same engine. On the MLX engine, oMLX led at 38.0 while LM Studio's MLX mode managed 17.0. The practical lesson for a UK buyer is that the wrapper can matter as much as the engine, so benchmark your own model, quantization and workload on your own machine before standardising on one.
View the data behind this chart
| LM Studio | llama.cpp | oMLX | Ollama | |
|---|---|---|---|---|
| Effective tok/s | tok/s41.7 | tok/s41.4 | tok/s38 | tok/s26 |
Power, Noise and the TCO Gap Nobody's Publishing Yet
This is where honesty matters more than a confident-sounding number. There is no verified wattage, electricity-cost or fan-noise figure in the current data for either the Mac Studio or a comparable Nvidia GPU workstation running LLM inference continuously. Any TCO spreadsheet built on invented power figures is worse than no spreadsheet at all.
What UK teams can do instead is structure the comparison properly before they ask a vendor for numbers: purchase price, expected resale value, and running cost are three separate line items, and the running-cost line needs a vendor-quoted sustained power draw under LLM load — not an idle or peak spec-sheet figure — multiplied by your actual UK electricity tariff. Before committing budget, it's worth taking the time to review the latest AI server cost index for GPU-server cost baselines you can hold a Mac Studio quote against line by line.
Mac Studio vs GPU Server: When Local Actually Makes Sense
The case for a Mac Studio is narrow but real: a single operator running a large quantized model privately, with no requirement for concurrent users, batching or CUDA-specific tooling. The M5 Ultra's 256GB and 512GB tiers let one desktop machine hold models that would otherwise need several discrete GPU cards, and the published data shows even the older M2 Ultra generating 11.1 to 12.13 tok/s on 70B-class models, which is workable for one person reading and writing alongside it.
The case against it is everything that scales past one user. The RPC split-inference data shows generation slowing when you combine Mac and Nvidia hardware for more capacity, and the H100-versus-M2-Ultra comparison shows a server GPU generating roughly twice as many tokens per second on the same 70B model class. If your workload needs batching, fine-tuning, or multiple concurrent inference streams, run the comparison properly: compare self-hosting LLMs with cloud GPU costs and use our AI GPU calculator to size the alternative before you commit to either path.
Future-Proofing: What to Consider Before You Buy in 2026
This year's supply swings are the future-proofing lesson. In May 2026, MacRumors reported that Apple had cut the M3 Ultra Mac Studio to a single 96GB configuration as memory shortages worsened; the 25 August refresh brought 256GB and 512GB options back with the M5 Ultra. Memory tiers can disappear and return with supply, so if your plan depends on the 512GB M5 Ultra, which Apple says arrives in late October, treat availability as a risk to manage rather than a given.
For UK teams who expect their model needs to grow — larger context windows, bigger parameter counts, multi-user services — the CUDA path remains the more extensible one, simply because VRAM pooling, batching and fine-tuning tooling are built around it rather than retrofitted onto it. Before assuming the cheaper-looking Mac path stays cheaper once your requirements change, see our AI GPU server shootout for 2026.
Sources
Every figure in this article traces to the sources below.
- •Apple — Mac Studio technical specifications (UK): M5 Max and M5 Ultra memory and bandwidth
- •Apple Newsroom — Apple introduces new Mac Studio with M5 Max and M5 Ultra (25 Aug 2026)
- •MacRumors — Mac Studio roundup: launch date, 512GB timing and US pricing (Aug 2026)
- •MacRumors — Apple cuts more Mac Studio and Mac mini RAM options as memory shortage worsens (5 May 2026)
- •Apple Newsroom — Mac Studio with M3 Ultra launch: 512GB and 600B+ parameter claims (Mar 2025)
- •Apple Newsroom — Apple introduces M2 Ultra: 800GB/s memory bandwidth (Jun 2023)
- •GitHub — llama.cpp distributed inference benchmarks: M2 Ultra Mac Studio + DGX Spark over 10GbE (Mar 2026)
- •GitHub — GPU benchmarks on LLM inference: H100 and M2 Ultra on Llama 3 70B (May 2024 snapshot)
- •GitHub — Performance of llama.cpp on Apple Silicon M-series: Llama 2 7B results table
- •Famstack — GGUF vs MLX on Apple Silicon, part 2: runtime comparison (Mar 2026)
- •eCorpIT — Local LLM on Apple Silicon vs cloud API break-even: M3 Ultra bandwidth (Aug 2026)
View the data behind this chart
| M4 Max 40-core GPU | M3 Ultra 60-core GPU | M5 Max 40-core GPU | |
|---|---|---|---|
| Generation tok/s | tok/s83.06 | tok/s88.4 | tok/s119.92 |
