I didn't say inference won't get cheaper, I said hardware is unlikely to for a long time. I do not believe there is a 1:1 correlation.
I've done work myself on the internals of AI models including mechanistic interpretability as well as adam-type training pipeline development and some other things I can't really talk about at the moment. Capability is going to continue to outpace hardware advances for a long time as long as the paradigm doesn't shift entirely to something like JEPA-type difference of approach, which could happen but isn't near-term.
But for the type of average mom & pop SMB pipeline, yes that won't require frontier level capability for much longer, if regulatory capture is held at bay which I am not convinced it will be, but am less certain about one way or the other.
For end users, the Billion+ people using GenAI will continue to rely on providers, and eventually they may fragment but there is a lot of incentive for Anthropic, OpenAI, Google, and Meta to continue to push hard to maintain the lead. Whether you or I can set up a pipeline is largely irrelevant for e.g. most of our relatives. From an ROI perspective, sure, which is why we'll hear even more how we can't afford to let anyone have access to Open models. We're fortunate that both Meta and Nvidia have signed on to that bandwagon, as much as I hate to admit it.
I don't expect the price of memory in Macs to drop for years, and in fact if I were to be forced to predict or wager, I'd lean more toward they may increase – which is why I think the 512GB Studio config option was delayed, not just due to availability. We'll see.