Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
P.S. I just noticed someone mention PrismML above. I too haven’t dug into the details, but suffice to say that it still looks more like proof of concept at the moment. They tackle fine-tuning the algorithm to work with highly quantized models that require less memory first, then figure out how to balance/optimize for precision/performance. The “easy” part is shrinking the model, the more difficult part is what is the limit before accuracy suffers to an unacceptable point. It also appears that their approach requires a lot more processing power to achieve similar results to existing quantized models, meaning at present their solution may overtax existing iPhone hardware and drain the battery at an unacceptable pace.
It's not just proof of concept.
You can download Bonsai Binary and Bonsai Ternary on your iPhone TODAY (eg via Locally AI) and experiment with them. Those two Bonsai models are based on Qwen3-8B, and they are pretty good (for 8B class models, of course!)
Neither do they require more power. Their white papers in fact give the required energy per token, and it's not at all bad (as common sense would suggest, given that moving data requires so much more energy than computation).
Screenshot 2026-07-09 at 1.33.01 PM.png



There's no obvious reason the same techniques would not work on either Qwen 2.6 27B (could shrink to about 4GB of total weights) or on eg a larger MoE model, fed into the device via Flash-MoE.

Or for that matter on Apple's existing AFM MoE model, quantized down to 1b from the 3 to 4b that Apple probably uses. (The paper describing the 3rd Gen AFM does not say what quantization they use, only that they use QAT, but based on the earlier 2nd gen model 3 or 4bits seems a reasonable guess?)
 
One county in Virginia has 37 data centers. I'm sure they didn't consider the implications of all those data centers when it came to their electricity bills. There is short term gain in jobs building the data centers.

And yet reality disagrees with your theory...
Screenshot 2026-07-09 at 1.37.54 PM.png


Want to try again, this time maybe telling us about some awful drought that has been killing people all across Loudoun county as they desperately struggle for unavailable water?
 
  • Like
Reactions: jjrtiger
Incredible the degree of extreme comments being made here by people who have no FSCKING CLUE what PrismML does, how the tech works, or why this is interesting and significant...

Here's a quick summary:
The point is not ONLY the intelligence density, though obviously that's nice; it's also that binary and ternary models can be executed with much lighter weight MACs (ie you can pack many more of them into the same area, or alternatively run them with substantially lower energy) than the INT8 or FP16 HW.

And Apple likely already has the designs for such binary/ternary optimal HW in house (having acquired it via their purchase of xnor.ai). It's generally believed (though I haven't seen anything absolutely definitive) that they are already using such a 1bit DNN and associated HW on AirPods for various tasks.

View attachment 2644307
PrismML can state whatever on a document. Actually executing their strategy is another story.
 
It's not just proof of concept.
You can download Bonsai Binary and Bonsai Ternary on your iPhone TODAY (eg via Locally AI) and experiment with them. Those two Bonsai models are based on Qwen3-8B, and they are pretty good (for 8B class models, of course!)
Neither do they require more power. Their white papers in fact give the required energy per token, and it's not at all bad (as common sense would suggest, given that moving data requires so much more energy than computation).
View attachment 2644310


There's no obvious reason the same techniques would not work on either Qwen 2.6 27B (could shrink to about 4GB of total weights) or on eg a larger MoE model, fed into the device via Flash-MoE.

Or for that matter on Apple's existing AFM MoE model, quantized down to 1b from the 3 to 4b that Apple probably uses. (The paper describing the 3rd Gen AFM does not say what quantization they use, only that they use QAT, but based on the earlier 2nd gen model 3 or 4bits seems a reasonable guess?)
Prism ML is just in the seeding stage of their company. It will be many years away until they can actually do anything and they will need way more capital investment. 16 million is not going to cut cheese.
 
👏 NOBODY 👏 WANTS 👏 THIS 👏 AI 👏 GARBAGE 👏
Honestly, I wish they would stop integrating this stuff so heavily into the OS (both Apple and Microsoft). Apple has been OK about this so far. But treat it like a feature and just tell us how much RAM we need to use it. If we don't have enough, then we just don't get the AI, and I think that's fine. Maybe RAM needs to be a "configurable" even on the phones where there are two options, AI and no AI. (Or you can choose the AI option and still turn off the AI if you just want more RAM.) I know that will never happen. But I'd prefer it be that way.
 
Last edited:
I’d love if they would also bake a bunch more not an AppleTV, or other home devices (Mac Mini) not just for when I’m on those devices but when my phone or other mobile devices could benefit from A.I. not everything needs to go to a data center, and *most* of the time I’m at home when I’m doing A.I. type stuff, even on my phone.

requests should be tiered and orchestrated based on the request.

A very simple conceptual example might be something like:

Orchestration/traffic-control/switch-board. Local device determines where and how a request is handled.

Simple stuff like “call my mom” or math, and other assistant type requests should mostly be handled on local device.

Slightly more complex, but still maybe a more personal level could be passed off to my household devices if they are available. (maybe even a new A.I. Hub device)

Truly global requests. and more higher order requests could be passed off to data centers.
 
  • Like
Reactions: HVDynamo
There’s 2.5 billion active users and subscribers across all AI services. That demand has driven up memory and storage costs. If the demand wasn’t there you would be paying 80 bucks for 32GB RAM sticks. But it’s 500 bucks instead.

Don’t be a tech denier otherwise you’ll always be posting cope material. Nobody gave you the right to speak for the rest of humanity and the economy.
Got any source for that 2.5 billion number you are throwing out? You know that's like 30% of the entire population of earth, right? Seems like a made up number to try to bolster your point. Also, it's a public forum, they have the right to post here just as much as you do.
 
  • Like
Reactions: turbineseaplane
These LLMs are power-hungry. The PC folks are running 'em on these 350-watt Nvidia GPUs. To do much on an iPhone you'll have to use the phone plugged in to AC power.
There are small language models as well and some really efficient ones that come from Chinese open source groups. They are not cutting edge but good enough for many uses.
 
It's not that people don't know what it is, it's that (to your point above) a zillion things have been roped into the same term of "AI".

A whole lot of it is just Machine Learning rebranded and there are excellent use cases and things we can get from that.
To add to this, it also has to do with how AI has been marketed and what it is seen as being used for. It's being shoved into everything, and many of those places just result in a worse user experience. There are absolutely good use cases for it and we should continue to develop it for those things. But I don't want AI touching everything I do. Google was better years ago than it is now, even with the AI search results. AI has made it worse. That's just one example.
 
  • Love
Reactions: turbineseaplane
Larger AI models take more ram than Apple puts on their iPhones.
I think it is the on device AI is not that much deal,it has much cons need big RAM and takes huge amount of storage, on devices and it is not more quicker than on older devices, like iPad m2 but I agree with on device ai for smaller features like text generation and some other smaller features. Cost of these devices will also be huge. and if Apple would want to do all ai features on device it would have this big cons, that’s my opinion.
 
You can’t run full 27 billion parameters on a 12-16GB RAM phone unless you reduce quality a lot to something like 1 bit or reduce active params.

There’s no magic sauce otherwise everyone would be doing it with PrismML.
Well, that's what this article is about. PrismML says they've figured out a major ingredient in the magic sauce of No-Loss Extreme Quantization, apparently before others have done so, or before others have announced having done so. And now Apple is seeing whether what PrismML has developed works well enough that Apple may want to buy the company or license their technology.

PrismML says it's figured out the math to achieve extreme 1-bit compression without ruining the AI's performance. They took the Qwen 27-billion parameter model, which is normally (currently) 54 gigabytes and requires a server rack to run, and compressed it down to just 4 gigabytes. The compressed model retains its complex reasoning and full "agentic" capabilities as if it were still sitting on a cloud server. If that's true, it could mean that it can run on a base iPhone, though maybe only on those with the upcoming 9 gig of RAM, and certainly on iPhones that contain 12 gig. Rumor was that Apple hoped to increase the RAM in the Pro models, starting with the 18, to 16 gig, and the base models to 12 gig, but the increase in RAM prices has put that off to some time in the future.
 
Last edited:
  • Like
Reactions: Tagbert
Larger AI models take more ram than Apple puts on their iPhones.
Well sure, full-blown current frontier-level AI in the Cloud requires more RAM than Apple is likely to put into iPhones for a good while, but that's not what Apple is aiming to do, at least not for a while.
 
I’d love if they would also bake a bunch more not an AppleTV, or other home devices (Mac Mini) not just for when I’m on those devices but when my phone or other mobile devices could benefit from A.I. not everything needs to go to a data center, and *most* of the time I’m at home when I’m doing A.I. type stuff, even on my phone.

requests should be tiered and orchestrated based on the request.

A very simple conceptual example might be something like:

Orchestration/traffic-control/switch-board. Local device determines where and how a request is handled.

Simple stuff like “call my mom” or math, and other assistant type requests should mostly be handled on local device.

Slightly more complex, but still maybe a more personal level could be passed off to my household devices if they are available. (maybe even a new A.I. Hub device)

Truly global requests. and more higher order requests could be passed off to data centers.
That's essentially how Apple is delegating these tasks currently, and I suspect it's also how it's done on Android phones running some form of Google AI.
 
In 10 years people are still gonna wonder when the bubble is gonna burst 😂

Everything better go on device because in 10 years all money ever in existence will have been sucked up at the current rates of spending growth.

It quite simply can not continue like it’s going right now.
 
There’s 2.5 billion active users and subscribers across all AI services. That demand has driven up memory and storage costs. If the demand wasn’t there you would be paying 80 bucks for 32GB RAM sticks. But it’s 500 bucks instead.

Don’t be a tech denier otherwise you’ll always be posting cope material. Nobody gave you the right to speak for the rest of humanity and the economy.
I’m fairly certain that number is only so high because phone assistants like siri and ai search results google is forcing at the tops of searches are included. 2.5B would literally be half of all adults on the planet and just under 1/3 of all people on the planet period. It’s also nearly 1/2 of *all* people who have internet access. I have doubts that that many people are actively seeking out the current ai choices beyond the phone assistants and search
 
It's not just proof of concept.
You can download Bonsai Binary and Bonsai Ternary on your iPhone TODAY (eg via Locally AI) and experiment with them. Those two Bonsai models are based on Qwen3-8B, and they are pretty good (for 8B class models, of course!)
Neither do they require more power. Their white papers in fact give the required energy per token, and it's not at all bad (as common sense would suggest, given that moving data requires so much more energy than computation).
View attachment 2644310


There's no obvious reason the same techniques would not work on either Qwen 2.6 27B (could shrink to about 4GB of total weights) or on eg a larger MoE model, fed into the device via Flash-MoE.

Or for that matter on Apple's existing AFM MoE model, quantized down to 1b from the 3 to 4b that Apple probably uses. (The paper describing the 3rd Gen AFM does not say what quantization they use, only that they use QAT, but based on the earlier 2nd gen model 3 or 4bits seems a reasonable guess?)

I may have played it a little loose with "proof of concept," but I can't just have you dropping truth bombs on me without some sort of informed response. So I read the white paper. lol

Ok, first, I want to say that what they claim (along with various benchmarks) is quite impressive. Their 1-bit Bonsai 8B model results sit right in the middle of the standard suite of benchmark test but use only 1/14th of the memory footprint. That's pretty impressive. But Intelligence Density is off the charts compared to equivalent parameter models. I will likely give them a spin when I have some time.

I think it still remains to be seen if this 1-bit inference acceleration approach scales. They are clear at the start of the white paper that their proprietary intellectual property is based on "mathematically grounded advances designed to preserve those properties [reasoning quality, behavioral stability] under aggressive compression." It will be interesting to see if their approach truly scales and reaches production with Apple. Their focus is on 8B and smaller models, but I'm curious if their approach will sale with large models as well given this approach is all about efficiency without sacrificing accuracy. For now I'll take their word for it, "Taken together, the Bonsai family illustrates that 1-bit design is not merely a compression technique, but a scalable approach to building practical, high-performance models across a range of deployment regimes. From the 1.7B model through the 8B model, Bonsai demonstrates that strong capability, efficient execution, and a smaller memory footprint can be achieved simultaneously rather than traded against each other."

And I did find on page 22 what probably caught my eye regarding my comment about power consumption. On an iPhone 17 Pro Max they observed 70% battery drain in 1 hour during sustained workloads. But given it's unlikely Apple or anyone with an iPhone would purposely run a sustained workload, this shouldn't be an issue. The fact is 1-bit Bonsai 8B model token generation far surpasses a 4-bit model on iPhone which results in much lower power consumption for practical inferencing purposes.

Power.jpg


All in all, I can see why Apple is interested. This is really a game-changer in terms of on-device/edge-device inferencing and should go a long way in easing local LLM memory requirements. The only thing that remains to be seen is if the performance and benchmark results truly and accurately deliver real-world results for tasks Apple expects Siri AI to execute.
 
Wowie wow guize!! look, more AI!!! :O ... can these people ever admit, even to themselves, that we've reached the much dreaded tech plateau and it's time to start lowering expectations in these tech companies? ... the party is over...
 
I find it a bit annoying that generative AI has taken over the term. ML and whatnot are, and have been, branches of AI, but now most people just use the term to refer to LLMs and diffusion models.

yes, I work in software, thank you. I was making this thing called a "joke"
I think you forgot your /s tags. We get posts like that daily, except that they have no idea about how things work, and are serious. They think that Apple can only work on one thing at a time.
 
Interesting approach. I dug deeper into this. Other models are trained with 16-bit weights, and then shrunk (quantized) by throwing data away to reduce the weights to 8-bit, or 4-bit, or 2-bit, etc. Once you get below 4-bit, most models hallucinate and make mistakes to the point of being unreliable or even unusable. PrismML flips this and trains the model to 1-bit weights from the start, which does yield much better results. PrismML is due to release their 27B model shortly, which supposedly can run on an iPhone 17 Pro (I would read between the lines and assume that means it was still too big to run on an iPhone 15 Pro/iPhone 16). We will have to wait and see how good this new model is. If they can prove that their training methodology scales with larger models, this would be great news.

Running a model on your phone requires more than just the core model weights - it also needs a KV cache that tracks what questions/information you sent to the model. For a long conversation, the KV cache may require as much memory as the model weights themselves. Fortunately, companies like Google have developed algorithms (TurboQuant) to compress the KV cache. So PrismML + TurboQuant drastically reduce the memory required to run a model.

However, to make a proper comparison, you must realize that frontier models (ChatGPT, Claude) are absolutely massive, using over 1T parameters (estimated, as they don't actually publish this info). Current open source models with 35B to 120B parameters are comparable, but definitely not as good.

So the real question that hasn't been answered yet is that even with all of these improvements in memory usage, will the model that you can run on your phone be good enough? My honest guess is "no" until phones start shipping with 16 to 32 GB of RAM.
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.