maccloud studio Reserve Now Portal
Introducing the maccloud mini

Your model box.
And your build box.

A dedicated Mac mini with the M5 Pro — 64GB unified memory, 4TB of storage and 10Gbps networking. It arrives with working models already loaded, and root access so you can replace every one of them.

Available now — no waitlist
Reserve — $350/mo Learn more ↓
18-core
CPU
20-core
GPU
64GB
Unified memory
M5 Pro chip
4TB storage
10Gbps network
Console access
Same idea as the maccloud studio, scaled down: one machine, one tenant, root access. The mini is the M5 Pro rather than the M5 Ultra — fewer cores and less memory, at a third of the price.
Pre-loaded, not locked down

Models are already on it. They're not the deal.

Your mini ships with a working local stack — a runtime, an OpenAI-compatible endpoint on the loopback, and a set of open-weight models sized to fit 64GB. Plug in and you're generating tokens the same day, without spending a weekend on build flags and quantization math.

That's a starting point, not a product boundary. You get root over SSH. Pull anything you want from Hugging Face, swap llama.cpp for MLX or vLLM, run whatever quantization you prefer, or delete our entire stack and start from an empty disk. Nothing we install is required to keep the machine running, and nothing phones home about what you run.

Open box, not a managed API. There's no allowlist of approved models, no per-token metering, and no proxy in front of your endpoint. The weights on the disk are yours to change, and the 4TB gives you room to keep a dozen of them on hand instead of re-downloading every time you switch.
Pre-loaded models on an open machine
Not just AI

A real CI host that happens to run models.

Most of what developers need a Mac in a rack for isn't inference — it's builds. The mini is a full macOS machine on a 10Gbps link, so point your existing pipeline at it: a self-hosted GitHub Actions runner, a GitLab runner, a Jenkins or Buildkite agent, Fastlane lanes, Xcode builds, simulator test matrices, notarization.

Because it's persistent and yours alone, the caches stay warm between runs. DerivedData, SPM and CocoaPods checkouts, node_modules, model weights — none of it evaporates when the job ends, which is where a lot of ephemeral-runner minutes go.

Root SSH — install your own toolchain
Out-of-band console access
Remote power reboot when a runner wedges
Switched PDU power control
The interesting part is combining the two. A pipeline stage that calls a model on the same machine — summarizing a diff, drafting release notes, triaging a failing test, scanning a log — runs over loopback. No per-token bill, no rate limit, no source code leaving the box to a third-party API.
Build, test, and a local model stage on one machine ONE MACHINE Build Test Local model No API egress
What fits in 64GB

Room for the models most teams actually use.

64GB of unified memory is shared between the OS, your build, and the model. These are the approximate weight sizes at commonly used quantizations — the ones we can load for you out of the box.

Model
Quantization
Weights
Llama 3.3 70B
Q4_K_M
≈40 GB
Qwen 2.5 Coder 32B
Q5_K_M
≈23 GB
Gemma 2 27B
Q5_K_M
≈19 GB
Mistral Small 24B
Q6_K
≈19 GB
Llama 3.1 8B
FP16
≈16 GB
Embedding + reranker pair
FP16
≈2 GB
Approximate weight sizes only — leave headroom for the KV cache, which grows with context length, and for whatever else the box is doing. Storage is 4TB, so keeping several models resident on disk is not the constraint; memory is.
Fit check

mini or Studio?

Take the mini if you…
Want a dedicated Mac for CI — Xcode builds, test runners, release automation — with a model on the side
Run models in the 8B–70B range and don't need to hold two large ones in memory at once
Need the machine now rather than at the end of a build window
Want root, console, and power control without paying Ultra prices for it
Take the Studio if you…
Need 256GB or 512GB of unified memory for frontier-size open models
Care most about tokens per second — the M5 Ultra's 1.2 TB/s is roughly four times the mini's bandwidth
Are serving inference to a team or an application rather than a handful of developers
Are fine-tuning rather than only running inference
Comparing the two? See the maccloud studio →
Configure

One machine. One price.

Available now — reserve and we'll provision it
maccloud mini · 64GB
M5 Pro · available now
$499/mo
$350/mo
or $17/day
ChipApple M5 Pro
CPU18-core
GPU20-core
Neural Engine16-core
Memory bandwidth307 GB/s
Unified memory64 GB
Storage4 TB SSD
Network10 Gbps
AccessRoot SSH · console · PDU
TenancySingle tenant, whole machine
Mac mini · 64GB · M5 Pro
Available now
$499/mo retail$350/mo launch price
or $17/day from account credit
Reserve and our team confirms your machine by email before any billing starts.
Daily rates are based on retail pricing and draw down prepaid account credit. Monthly and annual pre-pay are discounted.