Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
On the LLM angle: 16GB on an M4 base leaves you running maybe a Q4 7B (Llama 3.1 8B, Qwen 2.5 7B) at ~15-20 tok/s and effectively nothing else open — Safari with a few tabs plus the model process will start swapping the moment you also want to run your IDE. 24GB on the M4 Pro is where you can actually keep a 13-14B model comfortably resident and still work in a browser alongside it.

If the coding side is going to lean on cloud APIs (Claude, GPT-5, whatever's on your keychain), 16GB is genuinely fine and the extra cash is better in your bank account. If the point of the machine is local inference specifically, I'd wait a beat for the M5 mini — the neural-accelerator bump is meaningful for token throughput per watt, and 24GB on M5 will feel less tight than 24GB on M4 for the same workload.
 
I'd wait a beat for the M5 mini — the neural-accelerator bump is meaningful for token throughput per watt, and 24GB on M5 will feel less tight than 24GB on M4 for the same workload.
This makes alot of sense theoretically but the world is more chaotic than that.

We don't know when the 24GB Mac mini will be available, whether it will cost more and what the lead time will be.

If the OP can get their hands on an M4 24GB now then I would jump on it. When the M5 is released he can upgrade to it because used, the M4 24GB will likely sell for the same money as you can buy one or today.
 
  • Like
Reactions: Cape Dave
I now have the Mac mini in my possession and I’m very happy with how the final setup turned out.

What you’re looking at:

• Mac mini M4 – 512 GB / 16 GB RAM
• 2× Dell S2725QS – 4K / 120 Hz
• Edifier M60
• AJAZZ mechanical keyboard
IMG_1333.jpeg
 
@Archillias, congratulations on taking the plunge on the Mini M4 purchase. Also that is one sweet setup you have there, clean and minimalistic. Wishing you many enjoyable moments with the Mini.
 
  • Like
Reactions: SalisburySam
It’s been more than a year since the first reports on rumors sites that Apple was testing M5 versions of the mini and studio 😆
 
It’s been more than a year since the first reports on rumors sites that Apple was testing M5 versions of the mini and studio 😆
Yes. It's been a long wait. Apple cancelled my 64GB RAM M4 20Core GPU order last month. I was angry😡, but it was a blessing in disguise. In the past 2 weeks OpenAI dropped the API price of Luna by 90%. I use it for my business and I also have a ChatGPT Pro subscription for £89/month that I use for software engineering.

Basically, OpenAI has now made a great frontier model so cheap that it no longer makes economic sense for me to buy a Mac mini to run local models like GLM-4-9B, Qwen 3.5 or Gemma 3. 64GB Mac mini has been discontinued. For 64Gb I would have to buy a Mac Studio for and eye watering £3,500.

I use my ChatGPT Pro subscription heavily for coding. I haven't hit rate limits, can run 18 hour tasks and haven't exhausted my limits.

Therefore buying a Mac Studio 64GB to run local models would be like buying a bus to avoid a £100/month bus pass. 🤣

£3,500 in API provides 3.2 years worth of heavy AI usage. I might consider an M7 Pro if it provides a big leap in performance. But I'll have to wait and see what it performance is like compared to a ChatGPT subscription.
 
A bit of a weird update from an AI LLM running perspective.
The M5 Pro has a Single 16‑core Neural Engine and goes up to 64G of RAM.
The M6 basic has Dual 16‑core Neural Engine and goes up to a max of 32G of RAM.
So the M5 can run larger models but slower with MLX?
It'll be interesting to see which is better before purchasing.
 
Yes. It's been a long wait. Apple cancelled my 64GB RAM M4 20Core GPU order last month. I was angry😡, but it was a blessing in disguise. In the past 2 weeks OpenAI dropped the API price of Luna by 90%. I use it for my business and I also have a ChatGPT Pro subscription for £89/month that I use for software engineering.

Basically, OpenAI has now made a great frontier model so cheap that it no longer makes economic sense for me to buy a Mac mini to run local models like GLM-4-9B, Qwen 3.5 or Gemma 3. 64GB Mac mini has been discontinued. For 64Gb I would have to buy a Mac Studio for and eye watering £3,500.

I use my ChatGPT Pro subscription heavily for coding. I haven't hit rate limits, can run 18 hour tasks and haven't exhausted my limits.

Therefore buying a Mac Studio 64GB to run local models would be like buying a bus to avoid a £100/month bus pass. 🤣

£3,500 in API provides 3.2 years worth of heavy AI usage. I might consider an M7 Pro if it provides a big leap in performance. But I'll have to wait and see what it performance is like compared to a ChatGPT subscription.
I'm in the same boat. I bought the 64G Mini M4 Pro back when it launched. I can sell it now for more than I bought it!
Back then I wasn't sure if I would use cloud LLMs or local in the future.
So I splashed out on the Mini to future safe guard.
I'm now also an OpenAI Pro subscriber and it's doing all my work brilliantly for me.
I will add, however, that the Mini is being made to work by the Mac Codex App. It seems to use the local processing resources intensively for some tasks but I've been unable to find out exactly what from talking to it.
The Mini has proved to be a good investment so far (especially if I sell it right now!).
 
Last edited:
I'm in the same boat. I bought the 64G Mini M4 Pro back when it launched. I can sell it now for more than I bought it!
Back then I wasn't sure if I would use cloud LLMs or local in the future.
So I splashed out on the Mini to future safe guard.
I'm now also an OpenAI Pro subscriber and it's doing all my work brilliantly for me.
I will add, however, that the Mini is being made to work by the Mac Codex App. It seems to use the local processing resources intensively for some tasks but I've been unable to find out exactly what from talking to it.
The Mini has proved to be a good investment so far (especially if I sell it right now!).
Awesome that you managed to get your hands on one. Even a 16GB base Mini is sufficient for running cloud based models.

However if you already have a 64GB one then keep it unless you want to free up cash for a MacBook instead.

I have a 32GB M1 Pro MPB that I upgraded to from an M1 16 MPB 3 months ago. I can afford an M5 MacBook, however they are very expensive with high RAM and I would experience no difference in the way that I use my Mac for software engineering. M1 is still overpowered and I no longer do video editing or gaming on a mac to benefit from the latest chips. Running local models was the only area where I justified ordering new mac.

If the M7 brings big bumps in local dictation, improved spotlight(please) and other things, I might consider a Mini again. OpenAI have also recently announced their Jalapeño chip.

It enables them to serve models faster and cheaper. It will be rolled out later this year. Therefore next year, their models will be cheaper and faster to the extent that running open source wouldn’t make sense unless you want 100% privacy.
 
Awesome that you managed to get your hands on one. Even a 16GB base Mini is sufficient for running cloud based models.

However if you already have a 64GB one then keep it unless you want to free up cash for a MacBook instead.

I have a 32GB M1 Pro MPB that I upgraded to from an M1 16 MPB 3 months ago. I can afford an M5 MacBook, however they are very expensive with high RAM and I would experience no difference in the way that I use my Mac for software engineering. M1 is still overpowered and I no longer do video editing or gaming on a mac to benefit from the latest chips. Running local models was the only area where I justified ordering new mac.

If the M7 brings big bumps in local dictation, improved spotlight(please) and other things, I might consider a Mini again. OpenAI have also recently announced their Jalapeño chip.

It enables them to serve models faster and cheaper. It will be rolled out later this year. Therefore next year, their models will be cheaper and faster to the extent that running open source wouldn’t make sense unless you want 100% privacy.
I missed the Jalapeño chip announcement but I was reading today about NVidia's Groq3 LPX (Last year Nvidia paid $20 billion for Groq).
So it was no surprise to see the M6 boasting Dual 16 core NPUs but what that actually means still seems a bit vague until MLX benchmarks can explain it.
The NPU needs to evolve into a monster matrix multiplier way beyond what the GPU could ever do.
NVidia claim that with their chip a 5,000-token response that took 50 seconds now takes 1.5 seconds.
Just having a large memory isn't enough unless you can cascade matrix multiplication across is fast and efficiently.
Apple's silicon team are proven top players so I'm excited to see what an M7/8 can do on the desktop with running local LLMs.
These guys (OpenAI et al) are losing money hand over fist. Great that we're getting superb service at a reasonable price but at some point they need to put up their prices and squeeze us. It's at that point I'm hoping Apple's machines can become a more viable alternative.
 
  • Like
Reactions: Macalicious2011
I missed the Jalapeño chip announcement but I was reading today about NVidia's Groq3 LPX (Last year Nvidia paid $20 billion for Groq).
So it was no surprise to see the M6 boasting Dual 16 core NPUs but what that actually means still seems a bit vague until MLX benchmarks can explain it.
The NPU needs to evolve into a monster matrix multiplier way beyond what the GPU could ever do.
NVidia claim that with their chip a 5,000-token response that took 50 seconds now takes 1.5 seconds.
Just having a large memory isn't enough unless you can cascade matrix multiplication across is fast and efficiently.
Apple's silicon team are proven top players so I'm excited to see what an M7/8 can do on the desktop with running local LLMs.
These guys (OpenAI et al) are losing money hand over fist. Great that we're getting superb service at a reasonable price but at some point they need to put up their prices and squeeze us. It's at that point I'm hoping Apple's machines can become a more viable alternative.
I think next year we are going to see huge innovation in AI chips across Apple, Groq and just about any hardware manufacturer. Anthropic, OpenAI and lots of big corporations are developing their own chips. AI is currently CAPEX heavy because Nvidia has pretty much had a monopoly on chips.

However, in the past 12 months, the industry has realised that nvidia chips are optimised for training. It's better to develop seperate and cheaper chips that are cost and performance optimised for inference. Likewise for consumer/prosumer hardware, future devices could have chips optimised for inference especially if end users will not train their own models. This means more tokens per watt and GB of RAM.

For inference, the architecture of today's chips will likely be highly obsolete in 1-2 years now that chips designers know the exact use cases they need to optimise for. From Apple I expect less optimisation for memojis/image editing and more for local inference😛

Exciting times ahead. M7 is rumoured for first half of 2027 which isn't far away.
 
  • Like
Reactions: Seoras
I think next year we are going to see huge innovation in AI chips across Apple, Groq and just about any hardware manufacturer. Anthropic, OpenAI and lots of big corporations are developing their own chips. AI is currently CAPEX heavy because Nvidia has pretty much had a monopoly on chips.

However, in the past 12 months, the industry has realised that nvidia chips are optimised for training. It's better to develop seperate and cheaper chips that are cost and performance optimised for inference. Likewise for consumer/prosumer hardware, future devices could have chips optimised for inference especially if end users will not train their own models. This means more tokens per watt and GB of RAM.

For inference, the architecture of today's chips will likely be highly obsolete in 1-2 years now that chips designers know the exact use cases they need to optimise for. From Apple I expect less optimisation for memojis/image editing and more for local inference😛

Exciting times ahead. M7 is rumoured for first half of 2027 which isn't far away.
I think Apple should reconsider "NPU" and perhaps start referring to their "IPU" as that's really the top goal in silicon these days, as you pointed out, it's the inference that's the mass product not the training silicon.
Inference Processing Unit - IPU. Apple silicon boys are probably on it already let's see if marketing is too! 😉
 
I think Apple should reconsider "NPU" and perhaps start referring to their "IPU" as that's really the top goal in silicon these days, as you pointed out, it's the inference that's the mass product not the training silicon.
Inference Processing Unit - IPU. Apple silicon boys are probably on it already let's see if marketing is too! 😉
Good point. I think NPU will eventually be rebranded or become a chip of its own with cores for ray tracing, Face ID, Memoji, inference and other things. A bit like how Apple silicone has cores for different use cases like video encoding/decoding, security and graphics.
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.