deconstruct60
macrumors G5
It's a little hard for me to understand why Apple didn't realize (or just admit?) earlier that their PCC servers using M2 Ultra processors wouldn't be able to handle the workload for their customers once Siri 2.0/Apple Intelligence was fully up and running.
A central objective of Apples ML efforts for a long time were on how to compose the best model in the least amount of memory space. The PCC nodes were not really a move to throw memory control out the windows. The vast majority of Apple Intelligence instances would be on iPhones. The big move there was just to get to 8-12GB of RAM. So a M2 Ultra in a 4-way subnode would be a 4*64GB => 256GB model footprint. That is a order of magnitude larger footprint. Really should be able to do something better than what the local could do at that point.
It never was about running the largest possible'Frontier' model. It was about running something better than what you could run locally. It wasn't about wissing a biggest model pissing contest with OpenAI and Nvidia
Apple Intelligence even with PCC from the start had a 'punt to OpenAI" escape hatch. Similarly the AI summaries that Google is turning out for basic search and asnwer modest question is not the most computational resource intense model possible. Nobody is using 'super pro deluxe' models for that for the 'free' services. PCC never was for a 'for pay' AI workload model. It is for relatively modest scale like questions.
The real question at this point is why Apple would still be building M2 Ultra nodes. M3 Ultra maxes out at 96GB . The 4-way Thunderbolt 5 would make for a much better 'look no loose wires' connection if built a 4 Ultra logic board to mount 4 Ultras to. 4*96 384G is an even larger magnitude increase. M5 Ultra would be an even bigger easy cluster.
the M2 Ultra is from Jan 2023. Nvidia GraceHopper is from 2023-2024. It has fallen low enough getting onto the "Ok to sell to Chinese" list.
This looks a bit like the Mac Pro 2013 that had a 2nd GPU to handle workload and Apple somewhat fumbled getting buyin for OpenCL then switch focus to Metal and were limited software that really took advantage of the 2nd GPU computational horesepower. The hardware is there but the optimized , differentiating software stack to be layered on top didn't roll out the way they thought it would.
I think the 'moving goal post' problem that Apple has had is shifting desire to limit how much flows out the 'escape hatch' and whether an escape hatch only to OpenAI is a good idea or not.
The scaling requirements for the amount of off-device quality LLM inference Apple wants to provide weren’t totally anticipated even by researchers in the field in late 2022 when ChatGPT was released to the public, and maybe even by mid to late 2023 either, but by mid-2024 enough real-world user data had been collected though use of ChatGPT, Bard, etc. that it was becoming fairly apparent that the level of scaling needed to provide good responses for lots of users was far larger than anticipated.
M3 Ultra shipped late. On paper the stats would be better, but the hiccups through N3B likely meant there was little resolve to experiment with that. And if the Mac Pro wasn't going to help pay for M3 Ultra then already had cost control issues. Let alone PCC nodes which generate $0.00 in direct revenue.
And no planned M4 Ultra. ... so gap was going to be long until could get something more decent.
Should be at the point though were M2 Ultra should have been on verge of being phased out for M5 Ultra ( if skipping M4 Ultra was a very long term plan). Or phased out for a M4 Max. if a more limited memory footprint was projected. ( M4 Max is being droppedn by MBP 14"/16" so avaialble volume would be much higher. )
Among other things, it was found that training with a lot more data was needed to create larger, better models, but this led to increased requirements to provide quality inference compute after the training was done, and so some better idea of the hardware needed should have been more apparent.
The energy burn has not been on the same path as tthe model growth. The energy burn rate being higher is also a problem for Apple's "only renewable green datacenter' constraint.
But Apple began committing to using its M2 Ultra-based servers in mid-2024 anyway. Sure, Apple wanted its own hardware to do the job, but surely Apple should have known how to anticipate these things better.
M2 Ultra and associated RAM is likely easier to get at this point. Cheaper. Apple predict that OpenAI is going to go out and buy multiple billions of dollars of RAM with no product to sell in pure 'drunken sailor spend style'. Nobody was predicting that 2-3 years ago. Even OpenAI. Reactivating Three Mile Island wasn't a 'great idea' 3 years ago either. Now it is in the burn unlimited power dogma era.
The MBP Pro/Max 14/16 are rolling out in March 2026 instead of Oct/Nov 2025. Beat that wasn't on long term plan back in 2022-23 either. Previously ditto with the TSMC N3B major hiccups in the 2019 plan. Apple may be wealth, but it doesn't mean they control everything that might be a reqirement for product roll out. 'Stuff' happens.
Was Apple really thinking clearly when it assumed, even as late as mid-2024, that its iPhones and Macs would be powerful enough to do most of the work on-device, with cloud overflow being modest enough that data centers running its M2 Ultra servers could do the rest of the job?
Jumping from M2 Ultra to M4 Max or M5 Ultra would bring a larger memory multiuple to iPhones even with the recent increases to iPhone RAM.