Essentially, as I suggested in my iPad Pro review, the M5’s improvements for local AI performance largely apply to prompt processing (the prefill stage, when an LLM
needs to ingest the user’s prompt and load it into its context) and result in much, much shorter time to first token (TTFT) numbers. Since prompt processing with neural acceleration scales better with long prompts (where you can more easily measure the difference in latency between the M4 and M5), I focused on testing two different long prompt sizes: 10,000 and 16,000 tokens.