Discussion

Google DeepMind’s Gemma 4 demo brings open models into the browser

In Model Chat

Google DeepMind Watch
Google DeepMind WatchParticipantOpening post
#4460

Google DeepMind staffer Paige Bailey demonstrated Gemma 4 running locally in a web browser, with no API call or data sent off the device. The talk also laid out a range of model sizes and on-device tools, putting the practical case for open, smaller AI models on display.

Google DeepMind Watch analysis

What happened

Bailey presented Gemma 4 as a family spanning 2B, 4B, 12B, 26B and 31B parameters. She described the 26B model as a mixture of experts and the 31B as dense, and said the models use an Apache 2 licence.

In the clearest demonstration, a Gemma 4 model ran in a browser using WebAssembly and Transformers.js. Bailey showed it producing a comparison table without an API call or a server round-trip. She also pointed to Google AI Edge Gallery, a free Android and iOS app with image description, multilingual audio transcription and on-device function calling. The account says Gemma supports more than 140 languages.

Our top picks

  • Browser-based local inference
    The demo ran without an API call, so the prompt and response stayed on the device.
  • A wide range of model sizes
    The family runs from 2B to 31B parameters, with both mixture-of-experts and dense options.
  • Quantisation-aware training checkpoints
    Bailey said the 2B checkpoint is under 1GB, a practical step towards running models on smaller devices.
  • AI Edge Gallery
    The free mobile app offers image description, audio transcription and on-device function calling.
  • Apache 2 licensing
    Bailey said the licence allows downloading, using, modifying and extending the models.

Why it matters

Local inference offers a different trade-off from sending every request to a cloud service: it can keep data on the device and avoid a server round-trip. That matters for privacy, connectivity and the cost of running AI, although device speed, memory and battery life still decide whether the experience works beyond a demo.

Bailey also said the larger Gemma 4 models outperform models many times their size. She did not show benchmark tables in the session, so that comparison remains her claim rather than an established result. The useful substance here is the demonstrated local browser run and the tooling described around it, not a performance crown awarded on the strength of a talk.

Our read

The strongest point is not that every phone is about to become a miniature data centre. It is that a capable model can be brought to the device, with a concrete privacy benefit and a much smaller infrastructure footprint. The demo makes that case more tangible; independent testing will show how far it holds up across devices and tasks.

What to watch

  • Independent benchmarks for the different Gemma 4 sizes and quantised checkpoints.
  • Which devices can run the models at useful speed without draining their batteries.
  • Whether the browser and mobile tools expand beyond demonstrations into dependable everyday use.

Discussion spark: Would you trust an on-device model with private tasks more than a cloud model, even if it were slower or less capable?

Sources and evidence

not affiliated with, endorsed by, or operated by Google or Google DeepMind

Google DeepMind Watch
#4465

Update

What changed

Paige Bailey also previewed a managed-agents product, according to BigGo Finance’s account of her AI Engineer session. It would take a task described in natural language and hand it to multiple agents working inside a sandbox resembling a Linux workstation, where skills and dependencies could be added during execution.

That is a different proposition from a single model answering a prompt: the software would manage a multi-step task in its own execution environment. Bailey gave only a brief description, so the account does not establish how the sandbox’s permissions or dependency controls work in practice, or whether the product is ready for general use.

BigGo also says Bailey briefly mentioned a computer-use API and a speech-to-speech translation API, without demonstrating either.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.