Three Builds in 100 Minutes — DFlash Drops a Round-Trip, M1 Gets a 48-Hour Sweep, and Strix Halo Gets Static Rows
llama.cpp cut three builds in under 100 minutes — DFlash's speculative-decoding encoder fused into the KV-injection path, a 48-hour fa-vec sweep merged for M1, and RDNA3 mat-vec tuned for AMD's Strix Halo — while Qwen3.8-Flash-Next's official FP8 checkpoint keeps climbing and an open PR exposes a Vulkan TOP_K bottleneck costing 12 CPU round-trips per token past 1K context; HN's front page is now about whether agent harnesses are trustworthy.