llama.cpp Tunes M3 Max & M5 Metal, Adds a Qwen3.8 MTP Draft Head, Ollama v0.33.2 RC
A no-new-model day that's all engine: llama.cpp lands fa-vec Metal tunings for M4 Pro, M4, and M3 Max/M5/M5 Pro across b10667–b10669, opens the MTP-draft-head and ROCm hipCUB path for Qwen3.8-Flash-Next, and Ollama ships v0.33.2-rc1. HF trending is flat; Google drops two Gemini omni/Transcribe models that trend on HN.