Google DeepMind has announced Gemini 4 Argon, a model aimed at long software-engineering jobs and other demanding professional work. Its standout offer is the ability to generate up to one million output tokens in a single response, but access is still restricted and Google’s benchmark results remain company-run claims.
Google DeepMind Watch analysis
What happened
HackerNoon reports that Argon leads or ties on 14 of 19 benchmarks in Google’s comparison with GPT-6 Astra and two Claude models. On DeepSWE v1.1, a software-engineering benchmark, Google gives Argon a score of 77.9%, compared with 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. The picture is not a clean sweep: Astra leads on FrontierSWE v2 and an offline OSWorld-2.0 subset, while Opus leads on Terminal-bench 4.0 and PostTrainBench.
Google also lists introductory API pricing of $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%. It says those rates will later rise to $4 and $20 respectively. The same report says the model is currently available to selected cyber defenders through Google’s Fairwind programme, with paid API customers and Google AI Ultra subscribers next. Google has not given a broader release date.
Why it matters
Argon’s combination of long output, competitive reported results and lower listed token prices could make it worth testing for work that runs across many steps. The million-token ceiling is a capacity limit, not a promise that every task will benefit from a marathon answer; longer outputs can also raise the bill.
For now, most developers cannot check the model against their own workloads. Google’s benchmark table is useful for choosing what to test, but it is not an independent verdict, and the reported results vary by task.
Our read
This is a substantial model announcement with practical details, not just a new name on a leaderboard. The unusually large output window and published pricing make Argon one to watch, while the restricted rollout means any confident buying decision would be premature. Put the benchmarks on a shortlist, not on a pedestal.
What to watch
- When paid API access opens, and whether Google publishes a broader release date.
- How Argon performs on real software and professional workflows outside Google’s evaluations.
- Whether the introductory prices remain compelling once output length and total task cost are factored in.
Discussion spark: Would a one-million-token output limit change how you build long-running AI workflows, or is reliability on shorter tasks still the more important test?
Sources and evidence
- Gemini 4 Argon: How Does It Do Against GPT-6 Astra and Claude? – HackerNoon (7 October 2026, 00:19 UTC)
not affiliated with, endorsed by, or operated by Google or Google DeepMind