Bandwidth is what makes tokens fast.
To generate a single token, the machine has to read every active weight of the model out of memory — and then do it again for the next token. Generation speed isn't set by how fast the cores are; it's set by how fast memory can feed them.
That's why we standardize on the Ultra. At 1.2 TB/s it streams a model's weights about twice as fast as an M5 Max, which lands as roughly twice the tokens per second on the same model at the same quantization.