On the LLM angle: 16GB on an M4 base leaves you running maybe a Q4 7B (Llama 3.1 8B, Qwen 2.5 7B) at ~15-20 tok/s and effectively nothing else open — Safari with a few tabs plus the model process will start swapping the moment you also want to run your IDE. 24GB on the M4 Pro is where you can actually keep a 13-14B model comfortably resident and still work in a browser alongside it.
If the coding side is going to lean on cloud APIs (Claude, GPT-5, whatever's on your keychain), 16GB is genuinely fine and the extra cash is better in your bank account. If the point of the machine is local inference specifically, I'd wait a beat for the M5 mini — the neural-accelerator bump is meaningful for token throughput per watt, and 24GB on M5 will feel less tight than 24GB on M4 for the same workload.
If the coding side is going to lean on cloud APIs (Claude, GPT-5, whatever's on your keychain), 16GB is genuinely fine and the extra cash is better in your bank account. If the point of the machine is local inference specifically, I'd wait a beat for the M5 mini — the neural-accelerator bump is meaningful for token throughput per watt, and 24GB on M5 will feel less tight than 24GB on M4 for the same workload.