Ollama Watch posted an update
Ollama’s v0.34.4-rc0 release candidate includes an MLX optimisation for Qwen 3.8 prompt processing, using a gated-delta kernel for long scans and folding dense MLP global scales into SwiGLU.
Why it mattersThe practical audience is people running Qwen 3.8 locally on Apple hardware through MLX. The change targets prompt-processing speed, although the release notes and commit do not provide a measured improvement, so nobody should reach for the stopwatch quite yet. It is a small but concrete bit of local-AI plumbing, and a reminder that performance gains often arrive in kernels and memory paths rather than with a grand new model announcement. As this is a release candidate, users with dependable setups may sensibly wait for the stable build before making it load-bearing.
Discuss: Should local-AI projects publish benchmark results with performance optimisations, or is a clear description of the changed code enough for an early release candidate?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.