Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
To run a local AI for a specific task, you need 48GB or more to make it run comfortably. The models used are often around 32GB and when loaded into the unified memory can run very fast. It can be done via SSD as well, but that will slow the AI processing extremely, not to mention the wear it will cause on the SSD. There are smaller AI models for Mac's with less RAM, but at some point they are not as "smart", fast and useful as the larger ones. Of course it depends on the task at hand.

If more users do move away from AI in the cloud as the current trend shows, I'm sure that the demand for more memory in Mac mini's or Mac Studio's will increase big time and the worldwide shortage will continue for the coming years. Even when datacenters stop buying everything.
You’re overstating the RAM requirements quite a bit. Modern quantization and frameworks like MLX let you run really capable models in way less than 32GB. I run a solid 13-14B model at 4-bit quantization in just 10-12GB of unified memory and it’s plenty smart for coding, writing, and reasoning tasks. Even some MoE models around 20-30B effective size run comfortably on 24GB.

The idea that you need 48GB to “run comfortably” was maybe true a year and a half ago, but it’s not anymore. The smaller models aren’t nearly as dumb as you’re suggesting...the efficiency gains have been massive and continue to grow. A well-optimized 8-13B model today often outperforms a poorly quantized 70B from last year.

As for the SSD swapping argument, that’s also super exaggerated on Apple Silicon. The unified memory architecture and fast SSDs handle occasional swapping far better than people assume. It’s not ideal, but it’s not the performance killer it used to be.

The real trend isn’t “everyone needs way more RAM.” It’s that models keep getting smarter at smaller sizes. That’s where the progress is happening, not in just making everything bigger and hungrier for memory.

There are also APIs so you can use a hybrid model and save on tokens significantly. I use Grok/Claude & Llama 3.1

Llama 4 is going to be another big step.
 
I didn't know this. I'm still learning about AI tools, but my impression so far was that Intel was the preferred platform; you could run these models on Mac, but not as well.

For something like Stable Diffusion, what would you say is better, an M-series Mac, or an Intel PC?
Image generation probably benefits greatly from Nvidia GPUs more than anything else. My M5 Max can run ComfyUI but most of the forums are dominated by people running dual 3090 GPUs at the low end.
 
No, it hasn't. None of the "AI" companies are making a profit. None. It's a stock-pumping scam because the tech world is out of ideas.
Now that's a different subject entirely than what I was talking about and the Brooks quote you referenced. I'm not interested in the AI market, I am interested in AI technology. The underlying technology has grown and expanded wildly in the past 3 years, as evidenced by the fact that so many in this thread are reporting success and positive results from their locally-hosted AI tasks on apple silicon and other solutions like DGX and STRIX. That was a fever dream three years ago and now it's just a regular Monday.
 
Are you really believing this guy's claims? Sure, people have gotten some toy models to work at a snail's pace, but nobody is paying Apple's RAM prices for real LLM models to run on.
Enough of those that have, myself included. I bought one Mac mini M4pro with 48GB in January 2025, just before all the AI buzz started in the popular news channels and company ads. If the 64GB model was in my budget, I would have bought that one. No regrets what so ever - it speeds up the regular software I use (both work and private) and I can use LLM when I want to. I could have bought the standard model and replace the memory myself, but I barely have the time to do that.

Also... I would not say it runs on a snails pace. Yes, some of the cloud AI are indeed exceptionally fast, but you do pay premium for that too. For many tasks, it works fine. Connect the mini to a small battery + solar panel and you can run long AI tasks at zero energy costs 24/7.
 
Last edited:
  • Like
Reactions: p96 and EugW
Rhetorical flourishes and glittering generalities. We've been hearing some version of this for THREE YEARS now.
The CPU speed run from 6502/8080 to M4 went on for how many years? There were some missteps along the way like the G5 and the Itanic, or even P4 which turned out to be a dead end, but the general trend was clear.
 
How much memory are people generally getting for AI on those Mac minis? Are they mostly M4 Pros with 48 GB RAM or are many M4s with 24 GB or less? Just curious.
For LLMs that can help with writing or simple coding tasks: 24GB
For LLMs that can code ambitious features: 64GB
For LLMs that that can code whole applications: 128-512GB

Basically, the more vram, the better models you can run.

Memory and me,our bandwith are the biggest constraints. Hence why M2-M4 Ultra/M4 with 128-512GB are fetching big money in eBay even though M5 is the latest SoC.

For this use the Mac Mini and Studio have been exceptional value for money compared with windows alternatives. Mini and studio are also small and have very low power consumption compared with power hungry Nvidia chips.

We have now reached a point where it’s financially better to buy a $2-7k mini or studio than spending the equivalent on OpenAI and Anthropic every month.
 
Not to mention the privacy concerns that are resolved through running local resources.
+ knowing that no one can nerf the model! No rate limits or weekly quotas either. 😎

Some use a hybrid approach. Use cloud models to write plans and review code but use local models to do the crunch work in between.

Or use cloud models for very urgent tasks but local ones for exploratory or non time sensitive tasks.
 
The way the guy tells it you'd think Apple got on this bandwagon through proactive forward thinking designs decisions rather than sheer luck, problem is they are stuck with one fab for their chips, can't get enough capacity to satisfy demand and have a severe memory shortage problem that's not going away, how's that for forward proactive thinking.

It was proactive thinking. Apple started integrating Neural Engines into the A11 Bionic chip. That was released in 2017 and was clearly in design for a couple of years before it was released. Everything they learned from those chips went into the M-Series design. In other words, on-device AI has been in Apple's strategic plans for over 10 years. Did demand surprise them and outgrow their ability to meet it? Yes. Just like it for every other company in the industry. The gate for Apple at this point doesn't seem to be their ability to produce M-Series chips. It appears to be their ability to get memory. You can go on the Apple website and order any machine with any of their current M chips. You just can't get the memory configuration you might want. That's on the memory companies, not Apple.
 
It is quite insane what can be done with Gemma 4 and 16GB RAM (compared to what was possible for just 3 months ago) Things are moving fast and direction is clear: LOCAL llm is the future!
BTW: I recommend 24-32 GB as a minimum anyway.
 
Did demand surprise them and outgrow their ability to meet it? Yes.
100% Agree. 12 months ago the world of AI was different. A $200/month Claude code plan provided $5-8k worth of API usage. Basically, AI was heavily subsidised to the extent that it didn't make financial sense to use local models. Also, back then local models sucked ass.

However now, AI plans have usage based pricing, daily and weekly quotas. Also models like Gemma by Google are decent and Qwen 3.6 27B which takes up 30-32GB VRAM is really good for coding.

Put 2 things together and boom! Mac mini and Studio demand has sky rocketed.
 
  • Like
Reactions: p96
I run a DGX Spark here (128gb RAM) and an M5 Max (128gb RAM) and the Mac is faster, but the CUDA stack has better software support. There are some large, quality models that run well in MLX mode on the Mac, and I'm basically just waiting until I can buy whatever M5 Ultra is released, whenever it's released, specifically for local AI.
Given that you already have an M5 Max with 128GB of RAM, I would save that money for the M7 Ultra. The architecture is going to be noticeably more optimized for AI.
 
Given that you already have an M5 Max with 128GB of RAM, I would save that money for the M7 Ultra. The architecture is going to be noticeably more optimized for AI.
I think the M7 and M8 will be beasts.

The reasons are that Apple now has more data and information on the most popular AI use cases. Therefore they can optimise the chips and memory for inference rather than on-device Siri and..............Memojis 😛
 
For LLMs that can help with writing or simple coding tasks: 24GB
For LLMs that can code ambitious features: 64GB
For LLMs that that can code whole applications: 128-512GB

Basically, the more vram, the better models you can run.

Memory and me,our bandwith are the biggest constraints. Hence why M2-M4 Ultra/M4 with 128-512GB are fetching big money in eBay even though M5 is the latest SoC.

For this use the Mac Mini and Studio have been exceptional value for money compared with windows alternatives. Mini and studio are also small and have very low power consumption compared with power hungry Nvidia chips.

We have now reached a point where it’s financially better to buy a $2-7k mini or studio than spending the equivalent on OpenAI and Anthropic every month.

This is a great post, and contains a great example of why I turned "autocorrect" off. If I know what you meant, the keyboard should have too.

Especially with LLM powered autocorrect...

I actually have experimented with entirely local models that run on a 17 Pro and they can at least do a far better job of transcription, I'm honestly not sure if the LLMs really are good at autocorrect, since nobody makes an actual third party keyboard anymore, just plugins that take advantage of the feature. Apparently "new" autocorrect was based on GPT-2.5 and did seem better for a while.

Sorry to tangentially reply but I saw that typo and it triggered me.
 
  • Love
Reactions: Macalicious2011
I thought the whole point of including neural engines was to include cores that were designed just to run A.I. Even a lot of iPhones have those. What will future hardware have that current hardware is lacking?
 
Given that you already have an M5 Max with 128GB of RAM, I would save that money for the M7 Ultra. The architecture is going to be noticeably more optimized for AI.
I want more than 128gb. I'll buy whatever the next box is that I can buy with more than 128gb.

Also, the M5 I have now is my primary laptop and it's not available for AI workloads at times when I'm using it for other demanding tasks. And the fans are annoying so I have to wear my ANC headphones when I'm using it for AI tasks. So I'd love to add a dedicated resource along side of it, ideally in a desktop form factor with better cooling.

I'll buy that M7 Ultra too, whenever it exists.
 
  • Like
Reactions: Populus
I thought the whole point of including neural engines was to include cores that were designed just to run A.I. Even a lot of iPhones have those. What will future hardware have that current hardware is lacking?
AI is a general term for machine learning. For example some chips like Nvidia ones are general purpose. They are good for training models, generating text, generating images or video. However that's inefficient and a waste of resources if the main thing your users want to do is to generate text.

If you know that as Apple, then you can design chips that are more biased towards that, they might even run cooler or be cheaper to produce.
 
Yawn... What a missed opportunity with the MacPro (or even better, an updated Apple Silicon Xserve). Would be great to have tons of expandable internal storage for AI models (that a Mac Pro supported and Studio/Mini does not) ideally with pcie network cards with reasonable speeds (i.e. 100,200,400 gbps) that would be possible with the Mac Pro (and impossible with Studio/Mini) and ideally with
some extra expandable RAM (on top of the unified memory) that Apple could certainly have engineered for an updated Mac Pro.

Damn, I just don't understand why Apple killed the Mac Pro in the era of AI, when it finally could have made sense.
 
Are you really believing this guy's claims? Sure, people have gotten some toy models to work at a snail's pace, but nobody is paying Apple's RAM prices for real LLM models to run on.
You need to do your research. Go look at the many tech videos on youtube comparing Apple machines to other AI platforms and you'll see why Apple's unified memory strategy has paid off big time.

Also take a look at my signature below... 🙂 I bought and paid for the max 64G of memory when the M4 mini launched almost 2 years ago and I reckon I could sell it today for more than I paid for it then. This thing flies...
 
It is quite insane what can be done with Gemma 4 and 16GB RAM (compared to what was possible for just 3 months ago) Things are moving fast and direction is clear: LOCAL llm is the future!
BTW: I recommend 24-32 GB as a minimum anyway.
Yeah, I also think that the future of AI is local… but component prices go in the opposite direction. If we can’t afford a device with enough RAM…

But yeah, let’s hope the trend continues and the local models keep improving with less and less memory needed.

That or… a breakthrough innovation in the storage area with ultra-fast SSDs with bandwidth comparable to RAM memory (yes, what Octane was trying to be).
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.