Every configuration runs the same M5 Ultra — not Max, not a previous generation. Only unified memory changes between plans.
Why M5 Ultra?
Bandwidth is what makes tokens fast.
To generate a single token, the machine has to read every active weight of the model out of memory — and then do it again for the next token. Generation speed isn't set by how fast the cores are; it's set by how fast memory can feed them.
That's why we standardize on the Ultra. At 1.2 TB/s it streams a model's weights about twice as fast as an M5 Max, which lands as roughly twice the tokens per second on the same model at the same quantization.
Memory capacity decides what fits. Bandwidth decides how fast it runs.
Every maccloud plan runs the identical M5 Ultra — never Max, never a previous generation — so only the memory changes between plans. And every machine is single-tenant: the full chip, the full 1.2 TB/s, and the full Neural Engine are yours alone, with no neighbor stealing bandwidth mid-request.
Included with every unit
Real access. Not a black box.
You get root-level SSH, an out-of-band console, remote power reboot, and switched PDU control on every machine — manage it the way you'd manage your own rack, not the way a locked-down managed platform lets you.
SSH access
Out-of-band console access
Remote power reboot
Switched PDU power control
Fit check
Is maccloud right for you?
Good fit if you…
Run local LLMs (7B–400B+) or fine-tune your own models and need the full chip and memory bandwidth to yourself
Need more unified memory than any off-the-shelf cloud Mac offers — AWS's largest EC2 Mac instance tops out at 32GB
Want real SSH, console, and power control — not a locked-down managed API
Can plan around a 16–24 week launch window in exchange for locking in pre-launch pricing today
Probably not if you…
Need compute running today — the launch window won't fit an urgent deadline
Your workload is iOS/macOS CI, Xcode builds, or general Mac hosting rather than AI inference
Your models comfortably fit in 16–32GB and you don't need the extra headroom yet
You want a fully managed inference API with no server administration at all
Configure
Choose your memory.
Pre-launch pricing — reserve before general availability
ChipApple M5 Ultra
CPU36-core
GPU80-core
Neural Engine32-core
Memory bandwidth1.2 TB/s
Unified memory256 GB
Storage4 TB SSD
Network10 Gbps
AccessRoot SSH · console · PDU
TenancySingle tenant, whole machine
Mac Studio · 256GB · M5 Ultra
Estimated launch: 16–18 weeks
$1,499/mo retail$1,200/mo pre-launch
or $50/day from account credit
You won't be billed until your Mac Studio launches — reserving just locks in today's pre-launch price.
Daily rates are based on retail pricing and draw down prepaid account credit. Monthly and annual pre-pay are discounted.
Reserve
Mac Studio 256GB
$1,499/mo$1,200/mo pre-launch · Launches in 16–18 weeks
or $50/day from account credit
You won't be billed until your Mac Studio launches.
Request received
We've got your reservation for the 256GB configuration at $1,200/mo pre-launch pricing. Our team will follow up by email to confirm — you won't be billed until it launches.