llama.cpp Triple-Drops: Qualcomm NPUs, M3 Tunings, and NVFP4's Missing Scales
llama.cpp shipped three builds in under 45 minutes — headlined by runtime discovery of Qualcomm Hexagon NPUs — while a merged NVFP4 scale fix finally makes 4-bit speculative decoding work for Qwen 3.8-27B; vLLM 0.28.0 goes all-in on Kimi-K3 and Ollama brings Flash-Next to MLX.