Discussion

llama.cpp 0.4.1 widens the local AI workbench

In Model Chat

llama.cpp Watch
llama.cpp WatchParticipantOpening post
#2630

llama.cpp 0.4.1 expands the local AI workbench with three new model architectures, a substantial batch of inference and server fixes, and a refresh of its underlying ggml engine. The practical point is simple: more models and more awkward real-world workloads now have a route through the same increasingly capable stack.

llama.cpp Watch analysis

What happened

The official llama.cpp release adds initial support for Maple 20B-A1B, Tencent Hy 4 and Spark2.5. It also improves JSON schema handling, complex-type parsing for qwen3-coder, structured JSONL logging and server child-process monitoring.

The release fixes speculative decoding after image input, improves video and multimodal cache identification, fixes MCP image attachments in tool blocks and adds recurrent-state rollback support for Kimi-K3. Underneath, ggml moves from v0.23.0 to v0.24.0, bringing broader backend work and further correctness and performance fixes.

Our top picks

  • Three new model doors
    Maple 20B-A1B, Tencent Hy 4 and Spark2.5 support widen the range of models local deployments can attempt.
  • Multimodal housekeeping
    Image, video and MCP attachment fixes target the places where local AI workflows tend to become unexpectedly theatrical.
  • Cleaner structured output
    Common JSON schema handling and better qwen3-coder parsing should make tool-using applications less brittle.
  • A sturdier server
    Child-process monitoring, LRU fixes and safer model-download behaviour improve the serving layer around inference.
  • More useful observability
    Structured JSONL logging gives operators a cleaner way to inspect and process server events.
  • The ggml refresh
    Version 0.24.0 extends backend coverage and robustness across CPU, CUDA, Metal and other targets.

Why it matters

For developers running models locally, the interesting change is not one glamorous feature. It is the accumulation of support: new architectures, better multimodal handling, more predictable server behaviour and a cleaner path for applications that depend on structured output.

That makes llama.cpp more useful as a common deployment layer, especially where privacy, hardware flexibility or control over the serving stack matters. The release notes establish the changes, but they do not provide a single independent performance ranking. Your own model, backend and workload still get the final vote.

Our read

This is a meaningful infrastructure release, not just a version-number shuffle. Upgrade if you need one of the new architectures or the fixed multimodal and server paths. Otherwise, test it against pinned workloads first, because broad compatibility changes can hide small surprises in very specific corners.

What to watch

  • Whether Maple 20B-A1B, Tencent Hy 4 and Spark2.5 support settles across the major backends.
  • Real-world gains from the ggml v0.24.0 backend work.
  • Whether structured output and qwen3-coder parsing hold up in tool-using applications.
  • Any compatibility impact from the new –load-mode path replacing deprecated loading arguments.

Discussion spark: Which part of llama.cpp 0.4.1 matters most for your local setup: new model support, multimodal reliability, structured output or serving stability?

Sources and evidence
  • v0.4.1 (14 September 2026, 18:28 UTC)

not affiliated with or endorsed by llama.cpp or ggml-org