Unsloth’s latest release lets users train, test and serve decision models, while adding native support for ComfyUI models and reporting substantial QLoRA speed-ups on two hardware setups. It is a broad, practical update for people building and running open AI models, not just another item in a long changelog.
Watch Desk analysis
What happened
The GitHub release page lists version 0.1.904-beta, published on 7 October. Among its additions are training text or vision language models as decision models with QLoRA, testing them through the Decision API and exporting supported models to GGUF. The release also says supported decision models can be served through llama.cpp.
Unsloth says users can run supported ComfyUI image and video models directly from Hugging Face, with automatic recognition of ComfyUI checkpoints and local model folders. Its notes also report QLoRA training up to 4.1 times faster for Qwen3.5-35B-A3B and up to 3.3 times faster for Qwen3-30B-A3B on A100 and RTX PRO 6000 hardware. Those are the project’s stated results, not a general guarantee across hardware or workloads. Read the Unsloth release notes.
Our top picks
- Train decision models with QLoRA
Users can train text or vision models, test them through the Decision API and get confidence scores for each option. - Serve supported decision models locally
The release adds llama.cpp serving support, including for models that understand images. - Run ComfyUI models in Unsloth
Supported image and video models can run from Hugging Face, with local checkpoints recognised automatically. - Faster QLoRA for two named models
Unsloth reports up to 4.1x speed for Qwen3.5-35B-A3B and 3.3x for Qwen3-30B-A3B on the specified GPUs. - Sandbox model-run code
The release lists Bwrap for Linux, Seatbelt for Mac and MXC for Windows, with status and protection settings in Desktop. - Useful Desktop and training fixes
Users can ask about open browser pages, choose download locations and manage RAG embedding models from the RAG menu.
Why it matters
The update joins more of the local-model workflow in one place: training, testing, exporting and serving, alongside image and video model support. That could save practitioners from stitching together quite so many tools, while the platform-specific sandboxing gives them controls for code run by models.
The speed figures are worth noting, but they are scoped to named models and hardware. The release does not establish that every training job will see the same gains. For readers, the practical question is whether these features work well in their own setup, not whether a headline multiplier can do the benchmarking for them.
Our read
There is enough here to make this a meaningful release: decision-model workflows, broader ComfyUI support, training changes and desktop improvements all have concrete uses. The most appealing part is the breadth without requiring readers to pretend that every changelog entry is a breakthrough. If you use Unsloth, the decision-model guide and the exact hardware notes are sensible places to start; test the speed claims against your own workload before planning around them.
What to watch
- Whether Unsloth publishes reproducible details for the reported QLoRA speed-ups.
- Which ComfyUI models and formats prove reliable in everyday workflows.
- How the decision-model export and llama.cpp serving paths develop.
- Whether the new sandbox controls are available and effective across the listed platforms.
Discussion spark: Would you rather have training, inference and image-generation workflows brought together in one tool, or keep them separate for greater control over each stage?
Sources and evidence
- Trump Doesn't Understand AI, Says Godfather Of AI – NDTV (5 October 2026, 07:16 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.