Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
Honestly, this one will go on the history books.

The massive failure of Siri and the delays, ultimately resulting in Apple caving and doing business with direct competitor #1.

Jobs would've never allowed this. This was a clear mismanagement and lack of vision of Tim Cook and everyone involved with Siri. As an engineer and if I were to be an engineering manager at Apple, I would clean house entirely.
Just give it a rest, no one would know what Jobs would do in today’s business and tech world. Apple and Google have been in business together sharing info since 2002.
 
We, the Apple customer base, are all wearing brown shoes from walking in all of the barn yard stuff shoveled at us from every level of Apple's talking heads for the last three years. Sad to say there has not been a grain of truth in any shovel full of or mouthful of verbiage coming towards us.

Sad to say and see that so many of the Apple staff were conned into passing on the dis or mis information.

But that is now our reality everywhere we look in the USA. Trust no one's statements. Make no purchase decisions without LOTS of Due Diligence and even after that effort we can still get sucker punched.
 
  • Like
Reactions: teaneedz
Is this what my 16 neural cores are for on my Mac? To send my data to Google?
Unfortunately, yes. A MacBook M4 chip has an AI TOPS rating of about 40 from its neural engine. An RTX 5070 is about 1000 AI TOPS. An NVIDIA H100 is about 3000 AI TOPS. While the neural engine can identify text in a picture, it really is underpowered for running an advanced version of Siri.

For reference, you can run an average intelligence local LLM (9B parameter) at a decent speed on a 16GB RTX GPU, and this is far more advanced than what is available on your iPhone. You can run much better models on a M3 Ultra MacStudio, but that is still inferior to what Claude or ChatGPT are running on.

Until Apple (and TSMC) figures out how to shrink the power of a MacStudio into a phone or laptop, the only solution is to have it hosted remotely in some data center. Based on Moore's Law, this is probably 5+ years away.
 
Just give it a rest, no one would know what Jobs would do in today’s business and tech world. Apple and Google have been in business together sharing info since 2002.
Ah yes, the keyboard warriors that could manage it so much better than one of the most successful CEO's in history...
 
Later in 2025, Apple said that Siri would get an update at some point in 2026, though it did not provide a specific launch timeline.

Let me fix the article. In June 2024 Apple said an update Siri would come later in the year. Then incompetence took over
 
Honestly, this one will go on the history books.

The massive failure of Siri and the delays, ultimately resulting in Apple caving and doing business with direct competitor #1.

Jobs would've never allowed this. This was a clear mismanagement and lack of vision of Tim Cook and everyone involved with Siri. As an engineer and if I were to be an engineering manager at Apple, I would clean house entirely.
This. I hope Ternus is able to move Craig into some advisory role and bring in a new software leader. Today, I was playing around with Apple Intelligence (or whatever it actually is) and just kept asking, "why did they do it that way," "why can't I do this," "why?!" To add insult to injury, I had to work with Apple's broken iOS keyboard and battle its inherent bugginess. Siri remains the punchline of every joke. I hate to imagine Gemini powering Apple Intelligence but maybe the options are limited. I will be turning off Apple Intelligence completely before I allow Google to benefit though, regardless of what kind of frankensiristein this turns out to be -- even with private cloud compute and even if it means the keyboard still won't work unless gemisiri is enabled.
 
Oh, good…not.
Yeah it's that mix of...

- Cool, Siri will no longer be a complete moron (honestly this coulda been achieved even during pre-AI days and Apple sat on their hands. It's essentially rebadged 90's voice recognition (but not even that good as it tries to guess what you want with no basis to do so instead of just listening to your commands).

- How come Google's ads can now track my brain. Oh right...
 
Unfortunately, yes. A MacBook M4 chip has an AI TOPS rating of about 40 from its neural engine. An RTX 5070 is about 1000 AI TOPS. An NVIDIA H100 is about 3000 AI TOPS. While the neural engine can identify text in a picture, it really is underpowered for running an advanced version of Siri.

For reference, you can run an average intelligence local LLM (9B parameter) at a decent speed on a 16GB RTX GPU, and this is far more advanced than what is available on your iPhone. You can run much better models on a M3 Ultra MacStudio, but that is still inferior to what Claude or ChatGPT are running on.

Until Apple (and TSMC) figures out how to shrink the power of a MacStudio into a phone or laptop, the only solution is to have it hosted remotely in some data center. Based on Moore's Law, this is probably 5+ years away.
Apple will have to adopt the new wave of CiM ( Compute in Memory ) to really exploit AI on device. That involves shrinking the model with no loss in resolution to the point where the model sits permanently in memory and is computed in memory thus eliminating data transfer from CPU-GPU-NPU to memory and back again. Even Apple Silicon’s unified zero copy memory pool is not good enough. They’ll need to either embed NAND in the core of the NPU or GPU or put tensors in the memory itself. This will have the effect of speeding up operations due to zero data transfer and it will lower power draw significantly.

Something similar to what Anker just announced that they’ll be rolling out to all their sound products and headphones.

 
Honestly, this one will go on the history books.

The massive failure of Siri and the delays, ultimately resulting in Apple caving and doing business with direct competitor #1.

Jobs would've never allowed this. This was a clear mismanagement and lack of vision of Tim Cook and everyone involved with Siri. As an engineer and if I were to be an engineering manager at Apple, I would clean house entirely.

Massive failure implies you actually know their road map.

Given things like CloudKit, UIKit, etc... Apple has a knack for creating abstractions into other technologies. Personal assistants aren't the benchmark of AI just because they are what popularized AI. Apple has AI in many other places.

But given the pattern of CloudKit... do you want to join the arms race to try to come up with the biggest model, which so far, has been a money sink with seemingly zero actual result.

Or do you want to provide the means by which all models can be served to an entire ecosystem and capture a part of the token usage? Call 1 model to do OCR on a multi-lingual PDF locally to your laptop and pass de-identified data to a more complex model that is maybe a niche best in class medical comprehension model and then pipe that into a common language model like GPT for a human explanation.

I wouldn't call it a failure if Apple ends up being the first to democratize and get common usage of all these models on their billions of devices. Capturing 15-30% of that spend would be a smarter play than dropping $200 billion in R&D to get $5 billion in revenue.
 
  • Like
Reactions: Tagbert
There's a certain amount of "semantics playing" in the article when it says "It continues to be unclear if the new Siri and Gemini-powered Apple Intelligence features will use Private Cloud Compute or will run on Google's servers." - there's no reason why it couldn't be both simultaneously. Private Cloud Compute is a software stack. "Google Servers" can mean physical machines in a Google data centre.

Private Cloud Compute could run on Google servers.


100% PCC is hardware brother. Those are highly specialized servers that have 8 SoC's per 2U with a single power supply and have RDMA enabled so potentially 4 TB of RAM per server.

The software can glue it together. But PCC is Apples largest advantage in serving many models for their developers and creating an App Store of models, with model feature management, serving, tuning, privacy, etc all solved for...
 
It's pretty clear from the quote that they're confirming that all that's happening is Apple (with their help) is using Google's Vertex AI Studio, part of the Google Cloud Platform, to finetune Google base models into newer Apple Foundation Models.

They're already updated the on-device System Language Model in *OS 26.5, and they've started this week rolling out updated models in the Private Cloud Compute environment.

Separately, Apple has been an enterprise customer of Google Cloud Platform and Amazon Web Services for at least 5 years, if not more. Previously, AWS was their preferred cloud provider, and teams looking to deploy cloud-based infrastructure internally, would favor AWS.

The designation of Google as their preferred cloud provider likely only only means that teams are now being encouraged to deploy in GCP instead, which makes sense from a balancing perspective. Google's touting of that position also means that Apple is likely getting a significant discount on their cloud bills.

Pretty sure "preferred cloud provider" is just fluff. Apple has way way way more spend with AWS than Google. And has a large data center foot print as well.
 
100% PCC is hardware brother. Those are highly specialized servers that have 8 SoC's per 2U with a single power supply and have RDMA enabled so potentially 4 TB of RAM per server.

The software can glue it together. But PCC is Apples largest advantage in serving many models for their developers and creating an App Store of models, with model feature management, serving, tuning, privacy, etc all solved for...
No, not according to Apple. See the Apple’s article linked. Apple literally describe PCC itself as a software stack. There are hardware-specific nodes managing security, but these are separate from PCC itself. They act as gateways to the servers running PCC - they are not the servers themselves.

 
Last edited:
Apple will have to adopt the new wave of CiM ( Compute in Memory ) to really exploit AI on device. That involves shrinking the model with no loss in resolution to the point where the model sits permanently in memory and is computed in memory thus eliminating data transfer from CPU-GPU-NPU to memory and back again. Even Apple Silicon’s unified zero copy memory pool is not good enough. They’ll need to either embed NAND in the core of the NPU or GPU or put tensors in the memory itself. This will have the effect of speeding up operations due to zero data transfer and it will lower power draw significantly.

Something similar to what Anker just announced that they’ll be rolling out to all their sound products and headphones.


Apple's focus has been on software based optimizations. They have a paper called LLM in a flash. They have compressed existing models into very tight spaces and reduced latency. They can load 7b+ models into phone memory footprints.


And you're forgetting the GPU/CPU and TPU can already directly access the same memory objects, there's no copy need. And they put tensor cores in the GPU cores too in the M5. And their software abstracts the need to differentiate.
 
  • Like
Reactions: Tagbert
No, not according to Apple. See the Apple’s article linked. Apple literally describe PCC itself as a software stack. There are hardware-specific nodes managing security, but these are separate from PCC itself. They act as gateways to the servers running PCC - they are not the servers themselves.


How do you read that as a software stack when they literally are talking about how servers get high speed validation done in the supply chain, etc. They literally say the PCC node is a server.

But the non target ability and everything they talk about is only achievable on these nodes. There is certainly software on that. But the software will not run on any other node.

"

Introducing Private Cloud Compute nodes​

The root of trust for Private Cloud Compute is our compute node: custom-built server hardware that brings the power and security of Apple silicon to the data center, with the same hardware security technologies used in iPhone, including the Secure Enclave and Secure Boot.
"

Also, WSJ did a video interview and they literally pointed at servers and said these are PCC nodes.

It's also predicated on a Darwin based kernel as there are heavy kernel interactions.

But yes, PCC is 100% servers built and designed by Apple running in Apple data centers.
 
Its gotta be 'Private Cloud Compute' - can't see it any other way with Apple's privacy focus.
Google has an AI server setup that is a near twin of Apple's Private Cloud Compute. It does the same kind of things to prevent training and exfiltration of queries and data. It runs Google's AI model on custom processors designed specifically for LLMs. They may end up using that for Apple's AI. I think that would be a reasonable solution.
 
Google has an AI server setup that is a near twin of Apple's Private Cloud Compute. It does the same kind of things to prevent training and exfiltration of queries and data. It runs Google's AI model on custom processors designed specifically for LLMs. They may end up using that for Apple's AI. I think that would be a reasonable solution.


Not the same thing. The integration is much deeper as it involves the iCloud Keychain and the ability for Apple to run AI queries against *your* data with *your* crypto key. Becomes really powerful to be able to run AI compute operations right on the data instead of copying it in and out of data centers.

Apple also has a big efficiency advantage. They may be able to serve other peoples models cheaper than they can. It would be a key advantage for Apples hardware to be able to have context windows well beyond millions because they are not at the mercy of Nvidia GPU memory limits and power draws. RDMA in macOS 26.2 lets you get into the multi-terabyte context window.
 
How do you read that as a software stack when they literally are talking about how servers get high speed validation done in the supply chain, etc. They literally say the PCC node is a server.

But the non target ability and everything they talk about is only achievable on these nodes. There is certainly software on that. But the software will not run on any other node.

"

Introducing Private Cloud Compute nodes​

The root of trust for Private Cloud Compute is our compute node: custom-built server hardware that brings the power and security of Apple silicon to the data center, with the same hardware security technologies used in iPhone, including the Secure Enclave and Secure Boot.
"

Also, WSJ did a video interview and they literally pointed at servers and said these are PCC nodes.

It's also predicated on a Darwin based kernel as there are heavy kernel interactions.

But yes, PCC is 100% servers built and designed by Apple running in Apple data centers.
How do you read that as a software stack when they literally are talking about how servers get high speed validation done in the supply chain, etc. They literally say the PCC node is a server.

But the non target ability and everything they talk about is only achievable on these nodes. There is certainly software on that. But the software will not run on any other node.

"

Introducing Private Cloud Compute nodes​

The root of trust for Private Cloud Compute is our compute node: custom-built server hardware that brings the power and security of Apple silicon to the data center, with the same hardware security technologies used in iPhone, including the Secure Enclave and Secure Boot.
"

Also, WSJ did a video interview and they literally pointed at servers and said these are PCC nodes.

It's also predicated on a Darwin based kernel as there are heavy kernel interactions.

But yes, PCC is 100% servers built and designed by Apple running in Apple data centers.
The plan may have been to run the whole PCC syatem on Apple’s in-house servers. It might not have to be, and at this point, partnering with Google, it isn’t. Why all the delays, almost two year’s now. Apple along don’t have the data centre capacity. Data centres are expensive, and they’re not built quickly. Apple has invested about 18 billion into AI, that’s not nearly enough to cover the expense of building enough data centre capacity.
 

Attachments

  • IMG_0002.png
    IMG_0002.png
    30.2 KB · Views: 24
Last edited:
The plan may have been to run the whole PCC syatem on Apple’s in-house servers. It might not have to be, and at this point, partnering with Google, it isn’t. Why all the delays, almost two year’s now. Apple along don’t have the data centre capacity. Data centres are expensive, and they’re not built quickly.
They are also partnered with other vendors. I suspect what you will see is a marketplace of models on PCC nodes. I don't believe Apple ever took a serious run at a full LLM. Only LLM's that know how to ask other LLM's. There's a big difference between asking something about world knowledge vs asking something about your data or something on your laptop. Apple's biggest win is they figured out how to glue all those together transparently so you can switch from Gemini to Claude to whatever as a pulldown. And apple can monetize other peoples models.

Apple also has a massive memory density advantage that allows even things like Gemini to potentially run at trillions of parameters in PCC to start to do some really interesting things with personal context or large context tasking. A task you do for weeks or months rather than starting a new chat conversation because you ran out of session limit.

Using the M3 Ultra as an example, each SoC can have 512 GB of RAM. There's 8 in 2U, so that's a 4TB of RAM server. They can throw 25 of them in a rack pretty easily seeing as how they only have a single power supply, so we're talking 30kw racks with 100 TB of RAM. You would need 3-4 150 kw nvidia racks to deal with that context length.

Apple doesn't have the power of nvidia throughput wise, but they have memory density and power efficiency. They also have scale. Apple can spit out millions of these servers without much fuss.

BTW, Apple has had data centers in the millions of square feet for decades. They still primarily run their own data centers for iCloud.
 
I'm not a google fan, but gemini is really great. It's the best thing google has done in decades. I'm sure they'll find a way to make it worse soon.
 
Of all the top AI Gemini is at least better than ChatGPT but still garbage compared to Claude. At this point I'd rather just use my own LLM on a machine at home. Who needs AI anyway to turn your lights on and off. Will be disabled on my phone for sure
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.