Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
But what kind of AI would you run on it without CUDA? If you just get any old gaming PC with some nVidia GPU for a fraction of the price, you'll be running AI maybe a hundred times faster.

There was a time when Macs supported nVidia GPUs... There was also a time when external GPUs seemed to be a thing. Wouldn't it be better if you could just plug in an external nVidia GPU with CUDA and just run whatever you want on whatever you want orders of magnitude faster and for a fraction of the price?

The neural engine is great for Siri and genmoji I guess but you can't do any real work with it no matter how fast is, no matter how much RAM it has.
You’re quite mistaken.
Have you looked at the memory bandwidth of the M studio M5?
Clustering, I don’t think Thunderbolt is fast enough, Apple needs to come up with new fast connections.
But what kind of AI would you run on it without CUDA? If you just get any old gaming PC with some nVidia GPU for a fraction of the price, you'll be running AI maybe a hundred times faster.

There was a time when Macs supported nVidia GPUs... There was also a time when external GPUs seemed to be a thing. Wouldn't it be better if you could just plug in an external nVidia GPU with CUDA and just run whatever you want on whatever you want orders of magnitude faster and for a fraction of the price?

The neural engine is great for Siri and genmoji I guess but you can't do any real work with it no matter how fast is, no matter how much RAM it has.
you’re mistaken. M studio performs better than DGX spark. You should check few video benchmarks from Alec
 
I'm not talking about gaming, I'm talking about AI. A gaming PC just happens to have an nVidia GPU usually. AI can be limited by VRAM but as Macs have demonstrated, you can have a lot of VRAM and still be incredibly slow because AI workflows are optimized for CUDA and not Mac GPUs. Open ComfyUI and load up any video or image model and see for yourself: on a high end Mac you'll get maybe 1-5 images per minute while on even the slowest nVidia GPU you'll get 1 image per second. You can optimize for VRAM, with new workflows running under 6 or 8GB of VRAM, that's less of a problem. But if the GPU can't handle it because it's simply incompatible with the software, then you can have all the VRAM in the world and it still won't run, or it will run poorly. Speed and architecture cannot be replaced by just adding more VRAM, you need both, with VRAM being the lesser of the two in order of importance.

You are talking about a specific task and software optimized for CUDA. Most people are interested in how fast and well a PC can handle big models locally, and you cannot handle them with NVIDIA GPUs because of VRAM limitation.
 
But what kind of AI would you run on it without CUDA? If you just get any old gaming PC with some nVidia GPU for a fraction of the price, you'll be running AI maybe a hundred times faster.

There was a time when Macs supported nVidia GPUs... There was also a time when external GPUs seemed to be a thing. Wouldn't it be better if you could just plug in an external nVidia GPU with CUDA and just run whatever you want on whatever you want orders of magnitude faster and for a fraction of the price?

The neural engine is great for Siri and genmoji I guess but you can't do any real work with it no matter how fast is, no matter how much RAM it has.
I dont think this is accurate. Cuda is very helpful but it's not the only way to process the data. You can run many models on Macs currently and some optimizations make them run very quick. Cuda is king but it's not the only option.
 
5090 is ~1.8 TB/s and the RTX PRO was somewhat affordable for what it was at $8,000-10,000 street. Not now; prices just doubled.

Also the toolchains are a lot better, MLX is nowhere near CUDA and the difference is enormous. Time to first token is a lot faster, mixed precision e.g. NVFP4 offer a lot of advantages.

What is pretty mediocre is the little spark desktop thing, that is like $4,000 and is slower than an M4 Max at a lot of stuff that isn't NV specific.

But a whole Mac will use less than the power of one card, so.
I guess the while point of this article is that you can cluster 4 Mac Studios together resulting in close to 2TB of unified memory and an overhead of ”only” 25%.

I don’t know so much about this, but my understanding was that you can’t really run a big model on a single GPU, and Nvidia limit clustering of the GPU’s beyond consumer gear? I also understood that the fabrik to cluster the really expensive GPU’s is so expensive that a setup like that is way out of reach for mere mortals?

I also understand most things are optimised for Cuda, because why not when that hardware is available. At the same time, isnt there a possibility that once cheaper hardware exists it could get some products to consider looking into optimising for it?
 
I guess the while point of this article is that you can cluster 4 Mac Studios together resulting in close to 2TB of unified memory and an overhead of ”only” 25%.

I don’t know so much about this, but my understanding was that you can’t really run a big model on a single GPU, and Nvidia limit clustering of the GPU’s beyond consumer gear? I also understood that the fabrik to cluster the really expensive GPU’s is so expensive that a setup like that is way out of reach for mere mortals?

I also understand most things are optimised for Cuda, because why not when that hardware is available. At the same time, isnt there a possibility that once cheaper hardware exists it could get some products to consider looking into optimising for it?
Right, for a very small lab of a couple people a cluster makes sense. When you get into having to serve more than a few users these clusters don't make a lot of sense anymore despite all the hype around them. They need faster interconnects and another generation of improvement to get there.

I do think M7 will be amazing. Apple just sunset AMX and is going all-in on SME (M6 won't even run AMX on its neural engines anymore) so there is definitely a focused push going on inside Apple which is good.

IF I were buying for myself only to use at home and wanted a ~10 year computer I think the M7 will be substantially better than the M5, but I personally can't wait for it right now given the market and what I need to get done.

You could e.g. run an EPYC system with a few RTX Pro cards in your basement but it would draw 2.5 KW or more and you'd still top out below 512GB of memory, so in that sense being able to run less quantized models there is good benefit to Apple, but the performance just isn't the same, and NVFP4 in particular has some pretty neat things as well as Sparse Attention techniques that are coming out every week it seems like that help you run things on lower amounts of memory.

Personally I run as much as I can at full precision BF16, sometimes when I run into issues even going higher than that which is VERY SLOW, and for CPU workloads a lot of the work I do needs FP64, and nobody is good at that which sucks. These 20% single core performance bumps are good but I really wish they'd do something different with the Ultra CPUs vs doubling the Max, there's a lot of room left here much like when we had HBM on consumer cards very briefly (Radeon VII etc.) and FP64 was not utterly crippled.
 
But what kind of AI would you run on it without CUDA? If you just get any old gaming PC with some nVidia GPU for a fraction of the price, you'll be running AI maybe a hundred times faster.

There was a time when Macs supported nVidia GPUs... There was also a time when external GPUs seemed to be a thing. Wouldn't it be better if you could just plug in an external nVidia GPU with CUDA and just run whatever you want on whatever you want orders of magnitude faster and for a fraction of the price?

The neural engine is great for Siri and genmoji I guess but you can't do any real work with it no matter how fast is, no matter how much RAM it has.
Yeah you run up to 32GB models on an 5090 maybe - and that’s it. What does that have to do with frontier models?
 
Was just talking about this yesterday, was hoping that they would put in four additional thunderbolt five ports that way more units could be clustered.

Because of scarcity of ram and availability of units, it might be more practical although more expensive to buy four Studio Pro Max units instead of maxing out one unit and waiting six months for it. 🙄
And I think I answered that you make no sense because they chain.
 
  • Like
Reactions: JosephAW
You’re quite mistaken.
Have you looked at the memory bandwidth of the M studio M5?
Clustering, I don’t think Thunderbolt is fast enough, Apple needs to come up with new fast connections.

you’re mistaken. M studio performs better than DGX spark. You should check few video benchmarks from Alec
Wait wait wait. So your argument to this is to compare with a mini-windows PC? We are not talking about those. A full on desktop with even a 5080 with limited VRAM still beats the previously maxed out Mac Studios in some AI scenarios just due to the NVIDIA GPU. That is what that user was stating. And a 5090 just makes it so much better.
 
  • Like
Reactions: Kriss_De_Valnor
I do think M7 will be amazing. Apple just sunset AMX and is going all-in on SME (M6 won't even run AMX on its neural engines anymore) so there is definitely a focused push going on inside Apple which is good.

The AMX is part of the CPU core complex. It never was attached to the NPU cores. So the M6 not running it there would be entirely consistent with the previous 5 implementations.

Apple didn't really allow developers direct access to the AMX. Except for a very narrow subgroup inside of Apple, every developer is suppose to call an Apple library to get "ML compute" type work done and that library takes data and returns answers. So making that library call SME or the NPU instead in some contexts isn't a big shift on the outside usage of the APIs provided.

There have been a few folks who hacked the AMX opcodes and directly tried invoking them, but that code may bust on M7 . Apple will just say 'told you not to use it'. Moving to SME is just better since the compiler will give more stable long term to more developers than the 'indiirectly invoke magic potion' set-up with AMX.
 
The AMX is part of the CPU core complex. It never was attached to the NPU cores. So the M6 not running it there would be entirely consistent with the previous 5 implementations.

Apple didn't really allow developers direct access to the AMX. Except for a very narrow subgroup inside of Apple, every developer is suppose to call an Apple library to get "ML compute" type work done and that library takes data and returns answers. So making that library call SME or the NPU instead in some contexts isn't a big shift on the outside usage of the APIs provided.

There have been a few folks who hacked the AMX opcodes and directly tried invoking them, but that code may bust on M7 . Apple will just say 'told you not to use it'. Moving to SME is just better since the compiler will give more stable long term to more developers than the 'indiirectly invoke magic potion' set-up with AMX.
Correct, from what I understand the Neural Engines on M6 just won't run AMX at all. Agree that SME is a step in the right direction - my point was that Apple is consolidating internally which is showing some momentum.
 
5090 is ~1.8 TB/s and the RTX PRO was somewhat affordable for what it was at $8,000-10,000 street. Not now; prices just doubled.

But a whole Mac will use less than the power of one card, so.

You're forgetting about unified RAM. That 5090 is capped at the VRAM on the card. These bad boys can access 512GB.

The power usage is thing. A 3x Mac Studio cluster can run off a single domestic outlet.
 
  • Like
Reactions: dogrivergrad68
Wait wait wait. So your argument to this is to compare with a mini-windows PC? We are not talking about those. A full on desktop with even a 5080 with limited VRAM still beats the previously maxed out Mac Studios in some AI scenarios just due to the NVIDIA GPU. That is what that user was stating. And a 5090 just makes it so much better.
The DGX is supposed to be NVidia’s entry into the “home AI appliance” market for those who want to run various open source LLMs on their own hardware for use with claude code, codex, or other AI based coding or workflow automation tools. That’s where the mini has found a niche because it is silent and draws very little power compared to other solutions.
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.