Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
making on device as a key focus is something, actually delivering it is another.

We will see how it goes.
Yeah the strategy makes sense.

Most people (and I include myself in this) will probably be asking fairly easy questions that a local LLM can handle ('When is my first appointment next week?' 'What's the weather?', 'Read out my messages etc.').

The same goes for local agents ('tell me when X emails me', 'alert me when I get any information from Amazon about that package that I ordered etc.').

Then any heavy lifting will go to the (Apple) cloud.

You'd expect there to be more complex models on the Mac of course (more NPUs, RAM etc.)

But as you say, can Apple implement this reliably any time soon? There is a risk that even if Apple sells the privacy angle, any on device models will massively lag the capabilities of frontier cloud models.

And will this be a solution that 3rd parties would want to use? I guess there will be a pretty cheap token cost (as in free on device and presumably very reasonable if it has to go to the (Apple) cloud).
Time for more RAM!
I sense a Tim Cook style upgrade coming, where we will finally be able to choose out of x, y, z RAM options that we want in the iPhone Pro / Ultra models - with the promise of being able to run more advanced local models (Siri Pro! Siri Ultra!).

And given the price of RAM - and given that it's Apple - we can expect any upgrades to be excruciatingly expensive.
 
  • Like
Reactions: rosuna
I was taught to use it or lose it
so I prefer using my internal BI*
it's freely available to humans

*brain intelligence


Smh I can’t do this I can’t read all these replies… and this is just page 1, but this topic really brings out the great comments in ppl…wow
Just in time for me to finally throw in the towel and hit the “off” switch for AI on all my Apple devices! I miss the in-person live WWDC but with all this AI I can’t imagine they’d be as lively in 2026.

Amen brutha (sista?) ummm the last line was super visual (wouldn’t be “as lively in 2026” ) that’s a mouthful lmao… very imagistic and calls to mind alll the ai characters we’ve been seen a lot lately…

If goog wasn’t behind this I’d be like yeah let’s all check it out!shoot I could wait ten + yrs for the no goog version smh
 
  • Haha
Reactions: madmin
apple keeps talking about AI yet they have yet to release anything related to it. apple = vaporware
Translation?
Photo search?
There's a model called AFM built into iOS (since at least iOS 26.0) that can be queried via <span>FoundationModel.chatSession()</span>

(And that's just language-based AI, there's vastly more in non-language based AI, including eg video compression and operation of the cellular modem.)

If you don't know what's going on in this space, that's on you...
 
Part of me wonders if this AI hysteria had never happened and Apple just introduced a revamped Siri with the exact same features, built on the exact same architecture, whether the discourse around it would be completely different.
 
  • Like
Reactions: tranceking26
making on device as a key focus is something, actually delivering it is another.

We will see how it goes.
FFS.
Go to App Store
Download an app called Locally AI.
Look in the prefs. It should immediately have access to a model called AFM (Apple Foundation Model). You can start chatting with it immediately.

You can compare it against other models that can also be downloaded to your phone. Interesting examples include
- Granite (IBM)
- Gemma (Google)
- Bonsai (ternary model, using only weights of 0, +1, -1)
- Qwen
In my experience they are all pretty much the same, none is obviously much better or worse than the others. However some have specific functionality, like image analysis, that makes them the only choice for a specific task (like "what is that tree called" followed by a camera shot). Supposedly some are "reasoning models" but I haven't found a task I care about on the phone that's stressed a reasoning model more than a "non-reasoning" model in terms of the solution.

It's often startling how similar are the answers all of them give. I suspect there's a LOT of distillation going on from a common larger model (legal/negotiated vs illegal/illicit, who knows).

A disappointment right now is that AFM can't do more than the other models in terms of controlling your phone. For iOS 26, AFM seems to exist primarily as a way to provide a small LLM that can handle language queries generated by other apps. (I gave a quick reference to the API for this in my previous post.) I expect this to change with iOS 27; an integration of the language capabilities of AFM, the agentic capabilities of Siri, and general phone app control (probably a new set of APIs, but built upon the puppeting APIs already present for accessibility).
 
  • Like
Reactions: microChasm
While I agree, it should be noted that Image Clean Up and Image Playground (both of which run on-device IIRC) cause significant battery drain when used.
Claims like this, while true for you, will likely very much evolve over time.

1. I suspect the first version of each of these new functionalities is first coded for GPU (perhaps even with some AMX assist) because that's easiest. The code may move to ANE, once there's some time to think about how to structure it appropriately for ANE as a target.

2. ANE evolves every year (even if this isn't clear to outsiders because a headline number, like TOPs, doesn't change). It may be that this functionality can and will move to say ANE of M4 and later, but not to your generation of hardware. And the same is true, of course even for GPU functionality; I suspect that tasks that can use NAX will use noticeably less battery than on say M4.

Ultimately this is like complaining about ray tracing performance. Yes, it's true that if you have older HW you have a far inferior ray tracing experience than newer HW, and that sucks for you. But it doesn't change the fact that things are much better on the newer HW.
 
Exactly this. I think Apple needs to do a miracle if they want to run on device AI on the iPhones with 8 GB RAM. Probably most of the workloads will be on the cloud, like right now.
FFS.
These AI threads really shows you exactly how prevalent is Girardian mimicry...
So many people repeating opinions (mostly brewed up in Qatar and China) that they've been fed by the media, so little independent thought, so much confidence about "facts" that are UTTERLY wrong.

You do realize that LLMs run on phones RIGHT NOW? I have four installed on my iPhone. You have one installed on your iPhone as part of iOS26, even if you don't use it. Scroll up a few lines to my previous comment to see how to access it.
 
  • Like
Reactions: Ishimura and rosuna
Interesting to see how Apple balances the cost/availability?/power usage/capabilities of their hardware to do a good job of 'on device ai' vs their scaling back high memory hardware offerings...
 
I would like to believe that they do and probably more so than other large high tech entities, yet the AI wave is catching up with them and it might start biting them financially. Alphabet, Amazon, Meta and Microsoft have already pushed it quite far for both the environment and humanity, so currently Apple are just trailing behind.
Apple has a big leg up on the corporations listed above. And that leg is the Apple Watch. Sure Amazon MD upload a photo of your rash and get an instant diagnoses. Gonna put doctors out of business. /s
 
FFS.
These AI threads really shows you exactly how prevalent is Girardian mimicry...
So many people repeating opinions (mostly brewed up in Qatar and China) that they've been fed by the media, so little independent thought, so much confidence about "facts" that are UTTERLY wrong.

You do realize that LLMs run on phones RIGHT NOW? I have four installed on my iPhone. You have one installed on your iPhone as part of iOS26, even if you don't use it. Scroll up a few lines to my previous comment to see how to access it.
Sure they run on phones, but at what speed and quality compared to frontier models running in the cloud? Or something that's as good as, say, Qwen 3.6 27B without needing 128/256GB of memory. Would be happy to hear your thoughts on that.
 
apple keeps talking about AI yet they have yet to release anything related to it. apple = vaporware
Well, not entirely true

Sure, most of the stuff they initially announced back at WWDC 2024 is still yet to release, there still is some stuff out for it

Some is better than none, but still I get your sentiment
 
Well, here's to hoping

Only way to go from here is up after that blunder that was everything that was initially announced for Apple Intelligence in 2024
 
On device better mean no Internet required and it will do whatever you want none of this I'm not allowed BS...
Eh, I wouldn't hold your breath too much

Sure, on-device will be better because unlike the chatbots we've come to know from the likes of ChatGPT, Claude, Gemini, and others, you won't need an Internet connection for it, but I'm sure that just like those afforementioned chatbots, there will be some kind of guardrails built in

If they didn't, Apple would be facing similar issues that the chatbots have already because of them giving people advice it shouldn't have like on how to commit suicide and all
 
Always available, always private is a good feature. Also, that is free. Intelligence would not be comparable to frontier, but if Siri can call to their smarter colleague in cloud, that might solve some of it.

Pivotal question: How smart real time voice model one can actually host on a phone? If we can distillate GPT-4o level model there, this might be ok enough strategy. But that is a big if.

Trouble is that smartest models are multimodal and real time. One cannot match that experience by having a dumber model that needs to say "I do not know about that, just a bit, I'll call my colleague to check if he knows".

I have it hooked up to the action button on my iPhone, so I press that button (kinda the equivalent of pressing power to call up Siri) and LocallyAI pops up.

Model sizes tend to be around 4B. Bonsai is 8B, but that's a ternary model.
Multimodality exists, in the sense that a few of the models can "see/analyze" images you give them. (AFM cannot, even though it's supposedly a VLM model, ie a combined vision/language model; this is probably a limitation of the iOS26 APIs and I expect it to be fixed as of WWDC)
LocallyAI comes with its own speech recognizer which, honestly, I find substantially inferior to Apple's keyboard-built-in speech recognition. So that kinda sorta gets you some multi-modality.

What's very much missing (and this goes for pretty much every model everywhere) is a harness that understands that you're on a device with a screen, and that you're a human with a face and hands. What I want from multimodality is things like being able to say "You see that PDF I have open with Thai writing, can you translate it to English". That's not a "model capability" issue, it's a UI issue. Similarly while talking to the model I want to be able to open the camera and say something like "What's that building I'm pointing to"; again not so much a tech issue as a "putting the pieces together" issue.

My GUESS is that Apple understands this, and this is the sort wrapper "around" the model that we'll get in OS27 (well, we'll get the first part of it; like building the GUI in the 1980s, or building iOS in the late 2000s, for a few years each year will deliver a big new set of important features).
To see how limited things are (and most of this is not malice, it's just everything takes time and evolution) for example right now you can get Grok (or a few other large LLMs) as a CarPlay app. This can be very useful BUT there's apparently no entitlement available to allow access to GPS. So you can't make a query like "Hey Grok, tell me about this town I'm about to drive through" which could be really helpful! This is just oversight, and the sort of thing that tells you nothing about limitations, just that it takes a few years to polish anything!

As for
1. speed, I find that on my iPhone 15 Pro Max all the local LLMs generate at slower than I can read, but not unbearably slower. Prefill is also not a terrible wait, a few seconds or so. Generally the actual answer I want is within the first few sentences, then the LLM goes off on its usual trying to be helpful blather giving you more and more, so I just kill it at that point.
For comparison:
- Grok, which is obviously, like all the hyperscalers, serving many queries at once, does not generate a response much faster. CarPlay Grok, which "talks" its response is, IMHO, horribly slow; all the voices irritate the heck out of me because they are so slow. Presumably this will improve with time, but honestly I find a locallyAI response to be generated at about the same speed, but with less irritation (because of the difference in scanning text vs waiting for slow speech).
- at the other end, Google in basic "giving a pseudo-AI response on the google request web page" mode is absurdly fast. That's obviously some kinda hybrid, if you go into legit AI mode you drop down to Grok-like speeds. It's obviously where we'd all like to be, but no-one gets that yet [unless you have your own personal Cerebras maybe?!]

2. accuracy. CURRENT small models intrinsically hallucinate more, it goes with the territory. I don't think that's inevitable; my understanding is there are schemes for the net, as it generates a response, to track a "confidence level", and if those sorts of ideas work out, we could have a small model switch to RAG at that point rather than hallucinate a response. I've found that every "factual type" question I asked got a correct response (usually just a factoid I'd forgotten and couldn't immediately retrieve -- one of the hells of growing older :-( ) except for one technical question where all the small models clearly went off the rails. (It's unsurprising a small model wouldn't know a technical chemistry fact; problem, like I said, is that *current* models don't know that they don't know.)

A second element of accuracy is something like a translation task, and here I was very impressed. All the models did a much better job of translating technical Chinese than did both Apple's built-in translation and Google translate. Of course they were also a lot slower! Apple/Google translate are instantaneous, a local model takes 30 sec or a minute depending on the length of the passage. Essentially Apple/Google translate is like a 12 yr old, knows something about the languages but no domain expertise. The LLMs were like a somewhat knowledgeable 18-yr old. By comparison Claude was close to a student domain expert, not an actual chip designer, but say a 2nd year EE student.

So one way, for example, that local LLMs might sneak into our daily use is as translation apps move to more and more of an LLM-like design. (I assume both Apple and Google translate are built around something like BERT, but BERT is what, 2017 technology? So if they move to something like 2023 technology that would be nice bump in capabilities.)
 
FFS.
Go to App Store
Download an app called Locally AI.
Look in the prefs. It should immediately have access to a model called AFM (Apple Foundation Model). You can start chatting with it immediately.

You can compare it against other models that can also be downloaded to your phone. Interesting examples include
- Granite (IBM)
- Gemma (Google)
- Bonsai (ternary model, using only weights of 0, +1, -1)
- Qwen
In my experience they are all pretty much the same, none is obviously much better or worse than the others. However some have specific functionality, like image analysis, that makes them the only choice for a specific task (like "what is that tree called" followed by a camera shot). Supposedly some are "reasoning models" but I haven't found a task I care about on the phone that's stressed a reasoning model more than a "non-reasoning" model in terms of the solution.

It's often startling how similar are the answers all of them give. I suspect there's a LOT of distillation going on from a common larger model (legal/negotiated vs illegal/illicit, who knows).

A disappointment right now is that AFM can't do more than the other models in terms of controlling your phone. For iOS 26, AFM seems to exist primarily as a way to provide a small LLM that can handle language queries generated by other apps. (I gave a quick reference to the API for this in my previous post.) I expect this to change with iOS 27; an integration of the language capabilities of AFM, the agentic capabilities of Siri, and general phone app control (probably a new set of APIs, but built upon the puppeting APIs already present for accessibility).
Finally! I thought I was the only one using Locally AI. If nothing else, it gives a good preview of how far one can go with AI models with the limited resources of an iDevice.
 
  • Like
Reactions: rosuna
While I agree, it should be noted that Image Clean Up and Image Playground (both of which run on-device IIRC) cause significant battery drain when used.
Image Cleanup in its current form is an oxymoron, delivering results like soap did way back in 1998.
 
  • Like
Reactions: paczos
All of yall saying nobody wants this and that Apple is behind on AI are BOTH correct.

Why would the company prioritize something everyone hates and is also not profitable?

Maybe some of yall aren’t old enough to remember but there were a great number of mp3 players technically superior to iPod before it came out and a million smart phones before iPhone was released…
 
I can’t speak for other people but I have never sent ChatGPT a single prompt or inquiry. The only AI I have used is Google’s automatic AI responses they added to their search function.

I have never used AI to write a single thing or idea I have put out into the world.

I have never asked AI to edit anything I’ve put out into the world aside from extremely basic autocorrect such as “vlight” to “blight”.

AI has its uses in researching things (but it also hallucinates so if you know nothing about the topic it is still very dangerous) and I’m not completely anti-AI. But, I don’t care about this stuff, and I get that some people are eating it up.
“I’ve never ONCE used AI!

…I use it EVERY DAY! “

🥴🥴🥴
 
FFS.
Go to App Store
Download an app called Locally AI.
Look in the prefs. It should immediately have access to a model called AFM (Apple Foundation Model). You can start chatting with it immediately.

In my experience they are all pretty much the same, none is obviously much better or worse than the others. However some have specific functionality, like image analysis, that makes them the only choice for a specific task (like "what is that tree called" followed by a camera shot). Supposedly some are "reasoning models" but I haven't found a task I care about on the phone that's stressed a reasoning model more than a "non-reasoning" model in terms of the solution.
Locally AI is just a generic LLM app. The true power for small LLM lies in multi-adapter/agentic/tool-use design. The Apollo AI app from Liquid AI demonstrates that with built-in web search tool. And their model is just 1.2B. A lot of published research on Apple’s machine learning site suggests that trend.

The architecture for running AFM also shows much potential. That 3B AFM model alone is as fast as other 1B MLX models. No wait time for the model to load, probably compressed in RAM and run instantly. Some sources already claim it already uses mutil module adapters architecture for specialized task.

Which much new progresses have been made for the past 12 months on edge AI I’m genuinely exited for this year WWDC for what Apple would do with their on device AI.
 
  • Like
Reactions: name99 and rosuna
I couldn't care less about AI. I would much rather have Apple put their efforts into fixing bugs and promoting/rescuing Mac gaming (both native development and via emulation like proton/crossover).
 
  • Like
Reactions: tranceking26
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.