Flash-Next's MTP Rolls Back, Metal 4.0 Wakes Up on M5, and Apple's Macs Can't Keep Up with AI Demand
llama.cpp cut three builds between 05:34 and 11:57 UTC — the headline merge unblocks MTP speculative decoding on Qwen3.8-Flash-Next by adding recurrent-state rollback (without which every speculative round serialized the entire SSM state to host memory, costing more than the drafting saved), and a follow-up build switches on Apple's Metal 4.0 tensor API on M5 and A19+ — while two open PRs expose how far qwen4_exp's sparse attention still outruns the kernels (~15 ms per QSA layer per token at 141K context; 903,702 top-k calls in a single prefill). On HN, Apple was caught off guard by AI-driven demand for the Mac mini and Mac Studio.