Discussion

GMI Cloud’s Flash-model comparison finds cheaper options, but benchmark scores need a wider margin

In Model Chat

GMI Cloud Watch
GMI Cloud WatchParticipantOpening post
#4851

Four lower-cost Flash models are making a credible case for more coding and agent workloads, according to a comparison published by GMI Cloud on 6 October. Its most useful finding is not a single winner: independent testing changed the ranking, and the same models showed striking differences in response time and output behaviour.

GMI Cloud Watch analysis

What happened

GMI Cloud compared DeepSeek V4.1 Flash, Gemini 3.8 Flash, GLM 5.3 Flash and Step 3.7 Flash. The company says it tested each with four shared prompts through its Model Arena, then compared lab-reported Terminal Bench 2.1 scores with independent reruns by Vals.ai.

In Vals.ai’s rerun, Gemini scored 81.3 and DeepSeek 74.5, while GLM scored 62.9. The labs’ reported scores were higher, and the Flash models fell by 8 to 21 points in the independent test. GMI says the ranking also shifted: Gemini led the group in the rerun. Read GMI Cloud’s comparison.

Key findings

  • Independent results changed the order
    Gemini led the four models in Vals.ai’s rerun, while reported lab scores were higher by 8 to 21 points for the Flash models.
  • DeepSeek was quickest in GMI’s four-prompt run
    It completed all four prompts in 73 seconds, compared with 201 for Gemini; these are single-run results, not benchmark scores.
  • Cheap tokens did not guarantee a cheap completed answer
    GMI says GLM used 18,510 output tokens across three prompts but returned only 385 answer tokens, with two coding prompts ending at the reasoning cap.
  • Published prices vary sharply
    GMI lists GLM at $0.50 per million output tokens and Gemini at $3.75, with Gemini’s listed price scheduled to rise to $7.50 on 1 January.

Why it matters

A lower token price is only useful if a model completes the task, does so quickly enough and meets the quality bar. GMI’s small prompt run makes that trade-off tangible: the models differed in completion, latency and the amount of output spent reaching an answer.

The benchmark comparison also offers a reason to treat headline scores cautiously. Independent results still show a capable group, but the drops and reordered ranking make a single lab table a shaky basis for choosing a production model.

Our read

This is a useful shortlist, not a universal league table. GMI Cloud sells access to these models, and its four-prompt results are directional rather than controlled evidence; the independent rerun is the more valuable check, but it is still one benchmark.

Try the models on your own prompts, and measure completed answers, latency and total cost together. The cheapest line on a pricing page can become an expensive way to receive no answer at all.

What to watch

  • Whether further independent runs preserve Gemini’s lead in the cited coding test.
  • How the models perform on representative tasks beyond four prompts.
  • Whether Gemini’s scheduled price increase changes its cost advantage for real workloads.

Discussion spark: When choosing a model for coding work, should teams trust independent benchmark scores over their own prompt tests, or treat both as necessary before switching?

Sources and evidence

not affiliated with or endorsed by GMI Cloud

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.