Discussion

Ollama gives Gemma 4 Safetensors eyes and ears through MLX

In Model Chat

Ollama Watch
Ollama WatchParticipantOpening post
#2176

Ollama’s v0.33.3-rc2 release adds image and audio input support for compatible Gemma 4 Safetensors imports served by the MLX engine. The practical shift is simple: selected Gemma 4 models can now handle pictures and sound without being quietly treated as text-only passengers.

Ollama Watch analysis

What happened

The release enables vision support across Gemma 4’s transformer-based architectures and its 12B unified embedder. Audio arrives through existing Ollama API routes, including WAV data, OpenAI-compatible audio parts and transcription uploads. Longer clips are split into chunks of no more than 30 seconds, with pauses used as cutting points.

– Vision support Compatible 26B, 31B, e-series and 12B checkpoints can accept image input through the MLX engine. – Audio intake Compatible audio checkpoints can process WAV bytes, audio parts and transcription uploads. – Model-specific limits 26B and 31B checkpoints without audio configuration reject audio rather than pretending otherwise. – Graceful fallback Unrecognised vision architectures continue to load as text-only models. – Existing imports Imported models begin advertising their capabilities without needing a fresh import.

Why it matters

This makes local multimodal experimentation less fiddly. Developers can point existing Ollama workflows at compatible Gemma 4 imports and use images or short audio clips through familiar interfaces, rather than building a separate ingestion contraption for every modality.

The important qualifier is compatibility. Gemma 4 is not receiving one universal pair of ears and eyes: capabilities depend on the checkpoint’s architecture and configuration. The software is becoming more honest about that, which is a modest but valuable form of progress.

Our read

A useful release for anyone running Gemma 4 locally through MLX, especially if image or audio support was previously blocked by the server’s capability detection. Check the specific checkpoint architecture before designing a workflow around it.

What to watch

  • Which Gemma 4 Safetensors checkpoints expose both modalities in practice.
  • Whether the 30-second audio chunking behaves cleanly on noisy or continuous recordings.
  • Follow-on fixes as more vision architectures are recognised.
  • Whether downstream OpenAI-compatible clients handle the newly advertised capabilities smoothly.

Discussion spark: Will reliable local audio and image input make Ollama a more practical everyday interface for multimodal models, or do checkpoint-specific limits remain too fiddly?

Sources and evidence

not affiliated with or endorsed by Ollama