Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.

MacRumors

macrumors bot
Original poster


mac-studio-purple.jpg
Multiple new Mac Studio units can be clustered together over Thunderbolt 5 and RDMA (remote direct memory access), creating a shared memory pool for running larger AI models.

If four Mac Studio systems are clustered together, Apple says they can deliver 3x faster AI inference than running the same task on a single system, with the shared memory pool helping to load the largest and most demanding frontier-class open-weight models available today.

Article Link: New Mac Studio Can Be Clustered Together for Faster AI Performance
 
  • Angry
Reactions: Z-4195
But what kind of AI would you run on it without CUDA? If you just get any old gaming PC with some nVidia GPU for a fraction of the price, you'll be running AI maybe a hundred times faster.

There was a time when Macs supported nVidia GPUs... There was also a time when external GPUs seemed to be a thing. Wouldn't it be better if you could just plug in an external nVidia GPU with CUDA and just run whatever you want on whatever you want orders of magnitude faster and for a fraction of the price?

The neural engine is great for Siri and genmoji I guess but you can't do any real work with it no matter how fast is, no matter how much RAM it has.
 
The main requirement is Thunderbolt 5 right? So the Mac Mini with M5 Pro should be able to join in, along with MacBook Pros.

In a few years, maybe we can put together cheap clusters.
 
  • Like
Reactions: Equitek
But what kind of AI would you run on it without CUDA? If you just get any old gaming PC with some nVidia GPU for a fraction of the price, you'll be running AI maybe a hundred times faster.

There was a time when Macs supported nVidia GPUs... There was also a time when external GPUs seemed to be a thing. Wouldn't it be better if you could just plug in an external nVidia GPU with CUDA and just run whatever you want on whatever you want orders of magnitude faster and for a fraction of the price?

The neural engine is great for Siri and genmoji I guess but you can't do any real work with it no matter how fast is, no matter how much RAM it has.
1.2 TB/s of bandwidth is faster than any consumer GPU available today like 4090 etc. I don;t know what you've been drinking.
Prompt processing is also many times faster thanks to their compute improvements.
 
You can cluster multiple sizes together right? I got the 256 M5 ultra, not wanting to kill my budget further with the unpriced yet to be available 512gb. But I foresee a future where I may need more, perhaps 512 with an M7 Ultra in a few years. Would be nice to be able to cluster those together.
 
  • Like
Reactions: Equitek
But what kind of AI would you run on it without CUDA? If you just get any old gaming PC with some nVidia GPU for a fraction of the price, you'll be running AI maybe a hundred times faster.

There was a time when Macs supported nVidia GPUs... There was also a time when external GPUs seemed to be a thing. Wouldn't it be better if you could just plug in an external nVidia GPU with CUDA and just run whatever you want on whatever you want orders of magnitude faster and for a fraction of the price?

The neural engine is great for Siri and genmoji I guess but you can't do any real work with it no matter how fast is, no matter how much RAM it has.

Gaming and AI are 2 completely different world. Also, AI is mainly limited by the available VRAM, even the highest end 6090 will support 32GB Max, while the Apple architecture allows for up to 512GB (minus some for RAM use).
 
And all that for the low cost of a car.

Two 512gb m5 ultra will likely cost around 35k. And that would still not be enough to run the best open source models. For kimi k3 you would need at least 3 of those macs.
 
1.2 TB/s of bandwidth is faster than any consumer GPU available today like 4090 etc. I don;t know what you've been drinking.
Prompt processing is also many times faster thanks to their compute improvements.
5090 is ~1.8 TB/s and the RTX PRO was somewhat affordable for what it was at $8,000-10,000 street. Not now; prices just doubled.

Also the toolchains are a lot better, MLX is nowhere near CUDA and the difference is enormous. Time to first token is a lot faster, mixed precision e.g. NVFP4 offer a lot of advantages.

What is pretty mediocre is the little spark desktop thing, that is like $4,000 and is slower than an M4 Max at a lot of stuff that isn't NV specific.

But a whole Mac will use less than the power of one card, so.
 
Was just talking about this yesterday, was hoping that they would put in four additional thunderbolt five ports that way more units could be clustered.

Because of scarcity of ram and availability of units, it might be more practical although more expensive to buy four Studio Pro Max units instead of maxing out one unit and waiting six months for it. 🙄
 
A few notes that probably should have been mentioned in the article:

  • RDMA applications are not limited to AI models. RDMA is also useful in areas like scientific/grid computing, big data, etc.
  • RDMA on macOS was introduced in Tahoe 26.2 and should work on any Apple Silicon Mac with Thunderbolt 5. It's not limited to the 2026 Mac Studio.
  • You can cluster two, three, four, or even five Mac Studios together (though if you do five, you need to use a ring topology, which significantly hurts performance since it becomes impossible to directly connect each Mac Studio to every other Studio because of the number of available Thunderbolt ports). I'm not clear on whether you can do more than five machines; the developer documentation doesn't address this. I've also read that clustering more than two systems is buggy, but this will surely improve as time goes on.
RDMA is niche for now, but it's worth pointing out the very low power consumption on these systems. Running a cluster of Mac Studios could potentially save you big money over time when you consider electricity and cooling. We've come a long way since XGrid and the Virginia Tech System X supercomputing cluster.
 
Last edited:
Couldn’t the author go and describe that this could be done on an M3 ultra?

This isn’t the first article that is sparse.
 
And all that for the low cost of a car.

Two 512gb m5 ultra will likely cost around 35k. And that would still not be enough to run the best open source models. For kimi k3 you would need at least 3 of those macs.
To be fair, still looks like a bargain compared to the $55,000 Mac Pro 2019 +$6000 display and $1000 display stand.
 
But what kind of AI would you run on it without CUDA? If you just get any old gaming PC with some nVidia GPU for a fraction of the price, you'll be running AI maybe a hundred times faster.

There was a time when Macs supported nVidia GPUs... There was also a time when external GPUs seemed to be a thing. Wouldn't it be better if you could just plug in an external nVidia GPU with CUDA and just run whatever you want on whatever you want orders of magnitude faster and for a fraction of the price?

The neural engine is great for Siri and genmoji I guess but you can't do any real work with it no matter how fast is, no matter how much RAM it has.
VRAM is where it would help. You can have the fastest gpu in the world but if it has 8gb of RAM, it would be pretty much useless. More RAM means it can hold bigger models in it’s memory. Better models are all larger than what consumer level GPUs can hold.
 
1.2 TB/s of bandwidth is faster than any consumer GPU available today like 4090 etc. I don;t know what you've been drinking.
Prompt processing is also many times faster thanks to their compute improvements.
Speed is not the only thing that matters. Try running AI image or video models in ComfyUI on Mac and you will quickly find that they either don't run at all, they run on the CPU only, or they run on the GPU but an order of magnitude slower than what you'd expect simply because everything is optimized for CUDA and nothing is optimized for Mac. So you can chain any number of Mac Studios together and you still won't get anywhere near CUDA because the issue is not speed or performance, it's the ecosystem that nVidia has built out. The 4090 may be much slower on paper but it will actually run all the image models and it will do so pretty fast as opposed to the Mac. Every week there are optimizations coming out for CUDA, making generation twice as fast, 10 times as fast, etc... While the top end Macs are still far slower than even the slowest nVidia GPU – and this has nothing to do with RAM, power, or the GPU itself.
 
Gaming and AI are 2 completely different world. Also, AI is mainly limited by the available VRAM, even the highest end 6090 will support 32GB Max, while the Apple architecture allows for up to 512GB (minus some for RAM use).
I'm not talking about gaming, I'm talking about AI. A gaming PC just happens to have an nVidia GPU usually. AI can be limited by VRAM but as Macs have demonstrated, you can have a lot of VRAM and still be incredibly slow because AI workflows are optimized for CUDA and not Mac GPUs. Open ComfyUI and load up any video or image model and see for yourself: on a high end Mac you'll get maybe 1-5 images per minute while on even the slowest nVidia GPU you'll get 1 image per second. You can optimize for VRAM, with new workflows running under 6 or 8GB of VRAM, that's less of a problem. But if the GPU can't handle it because it's simply incompatible with the software, then you can have all the VRAM in the world and it still won't run, or it will run poorly. Speed and architecture cannot be replaced by just adding more VRAM, you need both, with VRAM being the lesser of the two in order of importance.
 
RDMA is niche for now, but it's worth pointing out the very low power consumption on these systems. Running a cluster of Mac Studios could potentially save you big money over time when you consider electricity and cooling.
This would be an interesting benchmark. A few months ago, I saw a post on some forum about a guy running a comparison of an image analysis pipeline. He ran it for 1000 images and compared a 5090 to a then current Mac Studio and tracked power usage and time. IIRC, the Mac had lower power usage but took about 10x longer to finish the task so total energy consumption was higher.

The takeaway for me at the time though was that he extrapolated this to 1,000,000 images and his estimate was that it would take around 20 kWh to complete.

It'd be interesting to see how this changes over time.
 
  • Like
Reactions: Soba
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.