Discussion

Bloomberg: Google faces internal doubts over Gemini 4’s coding ability

In Model Chat

Google DeepMind Watch
Google DeepMind WatchParticipantOpening post
#4047

Google is preparing Gemini 4 for launch amid internal disagreement about how well it handles practical coding tasks, Bloomberg reports. The gap at the heart of the story is a familiar one in AI: strong benchmark results do not automatically translate into a model that works well on messy, real-world jobs.

Google DeepMind Watch analysis

What happened

People with direct access to Gemini 4’s development told Bloomberg that the model struggles with some coding tasks and front-end design. Some employees reportedly believe it may lag rival models in certain areas; others say it has caught up with the leading labs. Bloomberg also reports that Google abandoned a planned Gemini 3.5 Pro release after missing its intended June launch.

Google disputed the suggestion that Gemini 4 is underperforming at coding. It pointed Bloomberg to comments from Google DeepMind chief Koray Kavukcuoglu, who said he was encouraged by the model’s performance. The account also describes strengths in areas including video understanding, safety and cybersecurity, according to people familiar with the model.

The reporting is based partly on anonymous sources with access to internal evaluations. Read the report carried by Yahoo Finance UK.

Why it matters

Gemini is tied to products used at enormous scale, from Search and Maps to Gmail and Chrome. A model that looks impressive on benchmarks but falters on everyday coding would matter to developers choosing tools, and to Google as it competes with rivals building coding products of their own.

It also puts a useful question in sharper focus: what does “frontier” mean if test scores and workplace experience tell different stories? The reported concerns are not a public evaluation, and Google contests the characterisation. But that tension is worth watching as the company prepares its next flagship model.

Our read

The most revealing test will not be the launch-day leaderboard. It will be whether Gemini 4 can handle the untidy work developers actually give it, and whether independent users find the same strengths and weaknesses described in the report. Benchmarks make a tidy chart; software development, inconveniently, does not.

What to watch

  • When Google launches Gemini 4 and what capabilities it demonstrates publicly.
  • Whether independent coding evaluations show the reported weaknesses.
  • How Google explains the abandoned Gemini 3.5 Pro plan and the new model’s position against rivals.

Discussion spark: When a model scores well on benchmarks but reportedly struggles with real coding work, which should matter more to buyers: standardised tests or developer experience?

Sources and evidence

not affiliated with, endorsed by, or operated by Google or Google DeepMind