Ollama’s v0.33.3-rc2 release adds image and audio input support for compatible Gemma 4 Safetensors imports served by the MLX engine. The practical shift is simple: selected Gemma 4 models can now handle pictures and sound without being quietly treated as text-only passengers.
Ollama Watch analysis
What happened
The release enables vision support across Gemma 4’s transformer-based architectures and its 12B unified embedder. Audio arrives through existing Ollama API routes, including WAV data, OpenAI-compatible audio parts and transcription uploads. Longer clips are split into chunks of no more than 30 seconds, with pauses used as cutting points.
– Vision support Compatible 26B, 31B, e-series and 12B checkpoints can accept image input through the MLX engine. – Audio intake Compatible audio checkpoints can process WAV bytes, audio parts and transcription uploads. – Model-specific limits 26B and 31B checkpoints without audio configuration reject audio rather than pretending otherwise. – Graceful fallback Unrecognised vision architectures continue to load as text-only models. – Existing imports Imported models begin advertising their capabilities without needing a fresh import.
Why it matters
This makes local multimodal experimentation less fiddly. Developers can point existing Ollama workflows at compatible Gemma 4 imports and use images or short audio clips through familiar interfaces, rather than building a separate ingestion contraption for every modality.
The important qualifier is compatibility. Gemma 4 is not receiving one universal pair of ears and eyes: capabilities depend on the checkpoint’s architecture and configuration. The software is becoming more honest about that, which is a modest but valuable form of progress.
Our read
A useful release for anyone running Gemma 4 locally through MLX, especially if image or audio support was previously blocked by the server’s capability detection. Check the specific checkpoint architecture before designing a workflow around it.
What to watch
- Which Gemma 4 Safetensors checkpoints expose both modalities in practice.
- Whether the 30-second audio chunking behaves cleanly on noisy or continuous recordings.
- Follow-on fixes as more vision architectures are recognised.
- Whether downstream OpenAI-compatible clients handle the newly advertised capabilities smoothly.
Discussion spark: Will reliable local audio and image input make Ollama a more practical everyday interface for multimodal models, or do checkpoint-specific limits remain too fiddly?
Sources and evidence
- v0.33.3-rc2: gemma4: image and audio input support (2 September 2026, 22:07 UTC)
not affiliated with or endorsed by Ollama