// the two machines, as their spec sheets say
Apple: "Combined with up to 512GB of unified memory and 1.2TB/s of memory bandwidth, 50 percent higher than before", and "Mac Studio with 512GB of unified memory is coming in late October." The specs page lists the base M5 Ultra at "96GB unified memory Configurable to: 256GB or 512GB (M5 Ultra with 36‑core CPU and 80‑core GPU)". NVIDIA: "With 32GB of GDDR7 memory, 1792 GB/sec of total memory bandwidth".
So "sixteen times" is the top configuration only (512 ÷ 32). The machine you can buy today starts at three times. And the prices are not comparable either way: the card's launch price buys a card that still needs a PC around it; the Mac's buys a whole computer.
Why memory speed sets the ceiling
When a model writes, each new token needs the model's weights moved from memory to the compute units. For one conversation on local hardware that movement, not the arithmetic, is the bottleneck. Apple's ML research team: "Generating subsequent tokens is bounded by memory bandwidth, rather than by compute ability." NVIDIA's inference guide calls it "a memory-bound operation". Databricks: inference "at smaller batch sizes—especially at decode time—is bottlenecked on how quickly we can load model parameters from the device memory to the compute units."
Two consequences. Memory your model doesn't fill adds no speed. And among boxes that already hold your model, the one with more GB/s has the higher ceiling. Apple's own measurement fits the rule: going from M4 to M5, bandwidth rose 28% ("120GB/s for the M4, 153GB/s for the M5") and writing got "19-27%" faster.
The speed limits, on one model
NVIDIA's own example model: "a model with 7 billion parameters … loaded in 16-bit precision … would take roughly 7B * sizeof(FP16) ~= 14 GB in memory." Divide each machine's GB/s by 14:
These are ceilings from a formula, not measurements. Nobody has published a measured M5 Ultra number yet, so this page does not claim one. NVIDIA lists DGX Spark at "128 GB LPDDR5x, coherent unified system memory" and "Memory Bandwidth 273 GB/s". RTX Spark PCs ("up to 128GB of unified memory", arriving October 2026) are left out because NVIDIA has not published their bandwidth.
Real Macs land under the line
llama.cpp keeps a community table of Apple Silicon results on the same LLaMA 7B model. Its 16-bit (F16) text-generation column is the 14 GB-class model the ceilings above use, so the two can be compared directly:
Speed rises with bandwidth, every chip lands under its limit, and the older multi-die Ultras land furthest below it. Rows used: the highest-core configuration of each chip (M1 Max 32-core GPU, M4 Max 40, M3 Ultra 80, M2 Ultra 76). The table has a 60-core M3 Ultra row too, and it is slightly faster in F16 (42.24). The file is 12.55 GiB (about 13.5 GB) rather than 14 GB; measured against the file size the shares are 67-83%.
Why "compare GB/s" is a rule of thumb, not a law. In the same table's 4-bit (Q4_0) column, the 614 GB/s M5 Max (119.92 tok/s) beats both 800 GB/s Ultras (94.27 and 92.14). A small bandwidth gap between different chip designs can flip. A big one, like the card's, is much harder to overturn. For scale on the card side, llama.cpp's CUDA thread records the RTX 5090 at 290-300 tok/s on that 4-bit file (3.56 GiB), a different model size from the table above, so it is not a point on it.
How to pick, and where this breaks
First, the box has to hold your biggest model. This is where the Mac wins outright: a model that doesn't fit on the card doesn't get the card's speed at all. A 70B at the common 4-bit format is large: llama.cpp's quantize README lists 70B at Q4_K_M as 43.1 GB, and Ollama's llama3.3:70b is 43GB (the plain q4_0 build is 40GB), both over a 32GB card. Unified memory is also shared with macOS, so no Mac loads a model the size of its full memory.
Then compare the GB/s line. Among machines that already hold your model, the higher number has the higher ceiling. Watch for Macs where the bigger memory option also changes the bus: the M6 Mac mini is 153 GB/s at 16GB and 170 GB/s at 24 or 32GB, and the Mac Studio's M5 Max is 460 GB/s at 36GB but 614 GB/s for the 48, 64 and 128GB configs. There the extra memory doesn't buy speed; the faster chip that comes with it does.
Related: are local LLMs worth it covers whether to run locally at all; quantization explained covers how a model's file size is set in the first place.
Sources: Apple Newsroom (Mac Studio with M5 Max and M5 Ultra; M6 and M5 Ultra), apple.com/mac-studio and /specs, apple.com/mac-mini/specs, Apple ML Research (Exploring LLMs with MLX on M5), NVIDIA GeForce RTX 50 announcement, NVIDIA technical blog (Mastering LLM Techniques: Inference Optimization), Databricks (LLM inference performance engineering), NVIDIA DGX Spark product page, NVIDIA blog (RTX Spark), llama.cpp discussions #4167 and #15013, llama.cpp quantize README, Ollama llama3.3 tags. Verified 2026-09-26. Speed limits and percentages are our arithmetic. Prices and availability are the vendors' and change.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
