Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
I use AI subscription (Claude code Pro, 20$ a month, 99% of time using Opus 5 with low effort mode to fit within limits) for programming apps, learning and research. I'm also playing with local AI by running models like Qwen 3.8 27B and recently Qwen 3.8 Next. I run them on M5 Max 128GB MacBook Pro 14 inch. My honest opinion after 2 months since I have the laptop and started testing different things is: local models are far far behind Opus 5 low effort in real day to day app development in agentic work. They are only good if you use them as a chat. This is mainly due to the hardware limitations and also partially due to the models a bit less capable than Opus 5.

As I'm building mobile apps with Opus, which is great for that I asked local models (Qwen 3.6 27B, Qwen 3.6 35B MOE, Qwen 3.8 27B and Qwen 3.8 Next) to implement the same app from scratch with the same prompt that I gave to the Opus as well.

Opus implemented fully functional app within 30 minutes, it worked and looked nice.

All those models I mentioned (except the Qwen 3.8 Next, for which 128 GB of memory is not enough because it needs a space for KV cache, so it failed at some point and I gave up) managed to generate an app and run it in simulator, however:
- it took many hours
- fan noise was huge
- app was not looking that good
- app was crashing
- app was not working as expceted

Probably if you spend additional time asking Owen to fix those issues for you, testing the app, it would be able to fix it. But after so many hours waiting and hearing noise I didn't wanted to continue. I don't want to work like that when I can have Opus 5 for 20$ a month.

For sure local models will improve and I see huge jump from Qwen 3.6 to Qwen 3.8. Also I see huge jump in speed thanks to the MTPLX where I can have 50 tokens/s in Qwen 3.8 Next.

On the other hand local AI is great for just chatting, asking questions, learning. it can explain things, you read it and ask another one. the context is small and it often managed to generate answer without running a fan at high speed.
 

Does it matter what the refresh rate is?

[TR]
[td]Support for up to five 4K displays, four 6K displays, or two 8K displays[/td][td]Support for up to eight 4K displays, eight 6K displays, or four 8K displays[/td]
[/TR]
 
I am very highly tempted to get the M5 Ultra so that I can run models locally and keep sensitive data on my own machine however I'm wondering how people who do this put guardrails in to stop agentic AI from running rampant with their local data, hacking their own network, etc.
It's all about isolation and permissions. The recent OpenAI debacle proved that agents will look for both strategic vulnerabilities and tactical security holes. Never in a million years would I give agents root access to a system, even if the system itself is isolated from others. Nor would I allow access to files, folders, or other personal information that I didn't want corrupted - largely due to the probabilistic nature of the thing itself. But building a strategic interface, giving only what it needs for assessment (and keeping source files away), and not giving open Internet access would be how I'd start.
 
This is the hair raising price in Canada. Although with 96gb ram, it's "not bad".
Now is just not the time to be buying high-end workstations unless absolutely necessary. The 64GB of RAM I purchased right before the datacenter mania now retails at 4x the price. I would gladly sell it if I had 16GB of spare DDR5 of any quality laying around
 
  • Like
Reactions: wlossw
Now is just not the time to be buying high-end workstations unless absolutely necessary. The 64GB of RAM I purchased right before the datacenter mania now retails at 4x the price. I would gladly sell it if I had 16GB of spare DDR5 of any quality laying around
The time isn't going to be right until the 2030s, if ever again.

There's too much to be gained from high unified memory and it's only going to improve with M7, new DGX, and beyond. That window where high performant large capacity personal computers were affordable is gone for a long, long time.

Even a 'bubble burst' will not correct all of this, over a billion people use GenAI and OpenAI just topped $1B in revenue from their ads service; this technology is not going away until something entirely replaces it, which may never happen. Sadly.

We are basically back into the 80s now where high performing computers cost the equivalent of a small older used vehicle. $10-20k is going to be the norm for a lot of professionals, and probably we'll see that creep up to $30k or even $50k for the very serious researchers (e.g. clustering 2 Studios for example).

The good thing is if you don't need a lot of memory or storage for your work, CPUs are really fast and efficient now. I feel for the DAW users who score film because those machines are now probably quadruple the price they were, and on the Mac side you can't go back to Intel since support is discontinued unless you pin the older OS, and then you still have crazy expensive storage and memory to worry about.
 
Last edited:
I use AI subscription (Claude code Pro, 20$ a month, 99% of time using Opus 5 with low effort mode to fit within limits) for programming apps, learning and research. I'm also playing with local AI by running models like Qwen 3.8 27B and recently Qwen 3.8 Next. I run them on M5 Max 128GB MacBook Pro 14 inch. My honest opinion after 2 months since I have the laptop and started testing different things is: local models are far far behind Opus 5 low effort in real day to day app development in agentic work. They are only good if you use them as a chat. This is mainly due to the hardware limitations and also partially due to the models a bit less capable than Opus 5.

As I'm building mobile apps with Opus, which is great for that I asked local models (Qwen 3.6 27B, Qwen 3.6 35B MOE, Qwen 3.8 27B and Qwen 3.8 Next) to implement the same app from scratch with the same prompt that I gave to the Opus as well.

Opus implemented fully functional app within 30 minutes, it worked and looked nice.

All those models I mentioned (except the Qwen 3.8 Next, for which 128 GB of memory is not enough because it needs a space for KV cache, so it failed at some point and I gave up) managed to generate an app and run it in simulator, however:
- it took many hours
- fan noise was huge
- app was not looking that good
- app was crashing
- app was not working as expceted

Probably if you spend additional time asking Owen to fix those issues for you, testing the app, it would be able to fix it. But after so many hours waiting and hearing noise I didn't wanted to continue. I don't want to work like that when I can have Opus 5 for 20$ a month.

For sure local models will improve and I see huge jump from Qwen 3.6 to Qwen 3.8. Also I see huge jump in speed thanks to the MTPLX where I can have 50 tokens/s in Qwen 3.8 Next.

On the other hand local AI is great for just chatting, asking questions, learning. it can explain things, you read it and ask another one. the context is small and it often managed to generate answer without running a fan at high speed.
How does that compare to Sonnet 5? (Can’t use Opus because the “financial controllers” say it is too expensive). Thanks!
 
M5 Max is the next step for me, as an upgrade from M1 Max 32 GB that's done fine since the pre-order arrived in 2022.

I opted for the 18-core CPU / 40-core GPU M5 Max, 128 GB RAM and 1 TB storage.

The M1 Max is still fast enough for a quantized Gemma-4-26b-a4b-qat model (with some parameter optimizations to take advantage of the M-series architecture model) but 32 GB of RAM leaves little room/available resources, KV cache, macOS, applications, and other user workloads simultaneously. Stepping up to the M5 Max:
  • About 1.5× the M1 Max’s memory bandwidth
  • 614 GB/s bandwidth
  • A substantially newer and faster GPU
  • Enough memory to run considerably larger models without swapping and increases limits on model size, context length, and multitasking capabilities.
  • Sufficient headroom for typical desktop use while inference is active
  • Much lower cost and power consumption than an Ultra
At >$5k USD, this will be the most expensive acquisition ever. I've been buying personal systems since C=64 and Compaq DeskPro 8086 days!
 
I don't work in AI, so I'm a little curious how one would connect them? Forgive my ignorance, but is it just the first one connects to the second, which then connects to the third, all in a straight line? Or does the last one also connect back to the first in the circle, or is it every Mac Studio connects to every other Mac Studio, so you need 4 cables per Mac Studio in a full mesh setup?
There’s one base computer and the other 3 connect to that base computer. So one computer has 3 cables connecting to 3 other computers. Those 3 computers only use that one cable connecting them to the main machine.
 
I would really like to understand more about this this "combining" works - couldn't afford to jump from 96 to 256GB of ram on the ultra. Could a max later be "linked" - it's not much more than the extra ram!
 
The time isn't going to be right until the 2030s, if ever again.

There's too much to be gained from high unified memory and it's only going to improve with M7, new DGX, and beyond. That window where high performant large capacity personal computers were affordable is gone for a long, long time.

Even a 'bubble burst' will not correct all of this, over a billion people use GenAI and OpenAI just topped $1B in revenue from their ads service; this technology is not going away until something entirely replaces it, which may never happen. Sadly.

We are basically back into the 80s now where high performing computers cost the equivalent of a small older used vehicle. $10-20k is going to be the norm for a lot of professionals, and probably we'll see that creep up to $30k or even $50k for the very serious researchers (e.g. clustering 2 Studios for example).

The good thing is if you don't need a lot of memory or storage for your work, CPUs are really fast and efficient now. I feel for the DAW users who score film because those machines are now probably quadruple the price they were, and on the Mac side you can't go back to Intel since support is discontinued unless you pin the older OS, and then you still have crazy expensive storage and memory to worry about.
I disagree with the headline prediction here on a number of levels. First of all, why should I believe there will be a shortage in compute long-term when there is no shortage of available inference at present prices? Inference has actually gotten way cheaper over the past six months not to mention people are getting smarter about using it. I design and build production AI pipelines for a living and the big two providers would gladly sell my company 100x above and beyond the tokens we actually need. When you combine this with the fact that chip manufacturers are rushing to fill demand I fail to see why I should expect prices to remain flat or even increase long-term.
 
Last edited:
I disagree with the headline prediction here on a number of levels. First of all, why should I believe there will be a shortage in compute long-term when there is no shortage of available inference at present prices? Inference has actually gotten way cheaper over the past six months not to mention people are getting smarter about using it. I design and build production AI pipelines for a living and the big two providers would gladly sell my company 100x above and beyond the tokens we actually need. When you combine this with the fact that chip manufacturers are rushing to fill demand I fail to see why I should expect prices to remain flat or even increase long-term.
I didn't say inference won't get cheaper, I said hardware is unlikely to for a long time. I do not believe there is a 1:1 correlation.

I've done work myself on the internals of AI models including mechanistic interpretability as well as adam-type training pipeline development and some other things I can't really talk about at the moment. Capability is going to continue to outpace hardware advances for a long time as long as the paradigm doesn't shift entirely to something like JEPA-type difference of approach, which could happen but isn't near-term.

But for the type of average mom & pop SMB pipeline, yes that won't require frontier level capability for much longer, if regulatory capture is held at bay which I am not convinced it will be, but am less certain about one way or the other.

For end users, the Billion+ people using GenAI will continue to rely on providers, and eventually they may fragment but there is a lot of incentive for Anthropic, OpenAI, Google, and Meta to continue to push hard to maintain the lead. Whether you or I can set up a pipeline is largely irrelevant for e.g. most of our relatives. From an ROI perspective, sure, which is why we'll hear even more how we can't afford to let anyone have access to Open models. We're fortunate that both Meta and Nvidia have signed on to that bandwagon, as much as I hate to admit it.

I don't expect the price of memory in Macs to drop for years, and in fact if I were to be forced to predict or wager, I'd lean more toward they may increase – which is why I think the 512GB Studio config option was delayed, not just due to availability. We'll see.
 
I didn't say inference won't get cheaper, I said hardware is unlikely to for a long time. I do not believe there is a 1:1 correlation.

I've done work myself on the internals of AI models including mechanistic interpretability as well as adam-type training pipeline development and some other things I can't really talk about at the moment. Capability is going to continue to outpace hardware advances for a long time as long as the paradigm doesn't shift entirely to something like JEPA-type difference of approach, which could happen but isn't near-term.

But for the type of average mom & pop SMB pipeline, yes that won't require frontier level capability for much longer, if regulatory capture is held at bay which I am not convinced it will be, but am less certain about one way or the other.

For end users, the Billion+ people using GenAI will continue to rely on providers, and eventually they may fragment but there is a lot of incentive for Anthropic, OpenAI, Google, and Meta to continue to push hard to maintain the lead. Whether you or I can set up a pipeline is largely irrelevant for e.g. most of our relatives. From an ROI perspective, sure, which is why we'll hear even more how we can't afford to let anyone have access to Open models. We're fortunate that both Meta and Nvidia have signed on to that bandwagon, as much as I hate to admit it.

I don't expect the price of memory in Macs to drop for years, and in fact if I were to be forced to predict or wager, I'd lean more toward they may increase – which is why I think the 512GB Studio config option was delayed, not just due to availability. We'll see.
I don't know your level of awareness of the business side of what's going on in AI currently, so I apologize if what I'm about to say is old news, but there are notable things happening that contradict this narrative. Meta and X are already selling their compute because they overbuilt capacity. Well, trying. Anthropic, OpenAI, and even Google cannot afford to sustain their most recently known capex. Nvidia is providing the lending for companies to buy their chips, because banks won't take chips as collateral. None of the hyperscalers have turned a net profit from AI. All signs point to decreased capex on chips while chip manufacturers are scaling up capacity.

You mentioned the 80s in your last post. Do you remember what happened between then and now? Moore's law mostly held true.
 
  • Like
Reactions: novagamer
Drooling over the M5 Ultra with 128 GB of RAM… for, eh… X-Plane 12.
But for that price might as well buy a Cessna 😁
 
I don't know your level of awareness of the business side of what's going on in AI currently, so I apologize if what I'm about to say is old news, but there are notable things happening that contradict this narrative. Meta and X are already selling their compute because they overbuilt capacity. Well, trying. Anthropic, OpenAI, and even Google cannot afford to sustain their most recently known capex. Nvidia is providing the lending for companies to buy their chips, because banks won't take chips as collateral. None of the hyperscalers have turned a net profit from AI. All signs point to decreased capex on chips while chip manufacturers are scaling up capacity.

You mentioned the 80s in your last post. Do you remember what happened between then and now? Moore's law mostly held true.
No apology needed, I don't doubt your pov at all, I actually agree with you – where we disagree is just on what the outcome will be regarding consumer hardware prices. I think you're on point otherwise, and you are more on the ground with business than me for sure, I'm much more heads-down in research this year in particular. Appreciate the civil discussion 🙂
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.