Every configuration runs the same M5 Ultra — not Max, not a previous generation. Only unified memory changes between plans.
Why M5 Ultra?
Bandwidth is what makes tokens fast.
To generate a single token, the machine has to read every active weight of the model out of memory — and then do it again for the next token. Generation speed isn't set by how fast the cores are; it's set by how fast memory can feed them.
That's why we standardize on the Ultra. At 1.2 TB/s it streams a model's weights about twice as fast as an M5 Max, which lands as roughly twice the tokens per second on the same model at the same quantization.
Memory capacity decides what fits. Bandwidth decides how fast it runs.
Every maccloud plan runs the identical M5 Ultra — never Max, never a previous generation — so only the memory changes between plans. And every machine is single-tenant: the full chip, the full 1.2 TB/s, and the full Neural Engine are yours alone, with no neighbor stealing bandwidth mid-request.
Included with every unit
Real access. Not a black box.
You get root-level SSH, an out-of-band console, remote power reboot, and switched PDU control on every machine — manage it the way you'd manage your own rack, not the way a locked-down managed platform lets you.