Discussion

Reflection’s first open model makes efficiency its pitch against China’s leaders

In Model Chat

Watch Desk
Watch DeskParticipantOpening post
#4904

Reflection AI has introduced Beam, its first open-weight model, and is pitching reasoning efficiency rather than a clean sweep of the model rankings. The reported results put it near an earlier GLM model on one coding test, while newer Chinese models score higher on several comparisons.

Watch Desk analysis

What happened

Reflection is a US AI startup founded by former Google DeepMind researchers. Its Beam model is a mixture-of-experts system with 501 billion parameters in total and 23 billion active per token, according to 36Kr. The publication says the company trained it on 23.8 trillion tokens and used 10,500 Nvidia GB300 GPUs for four weeks of reinforcement learning. Beam is in final red-team testing; Reflection plans to release its weights, technical report and development tools in October.

The company’s own results, as reported by 36Kr, put Beam close to GLM-5.2 on the DeepSWE v1.1 coding test, scoring 44.4 against 44.0. On that test, GLM-5.3 scored 61.0, Kimi K3 scored 68.0 and DeepSeek V4.1 Flash scored 74.2. Beam also scored below those newer models on Terminal Bench v2.1 and Humanity’s Last Exam. Reflection’s argument is that Beam can reach comparable performance on some advanced reasoning tasks with about a quarter to a third of GLM-5.2’s reasoning compute. Those are company-reported comparisons, not independent evaluations.

Why it matters

Beam is a substantial attempt to put a US-built open-weight model back in the contest, but its early case is about how much capability buyers get for their computing budget, not being the outright leader. That distinction matters to organisations weighing model quality against the cost of running it.

The reported training effort is also a reminder that “open” does not automatically mean cheap to build. Beam’s weights are not yet available, so users cannot test the model themselves or check whether the claimed efficiency holds on their own workloads.

Our read

The efficiency pitch is worth watching, particularly if the model’s weights arrive with enough detail for independent testing. But the rankings in this account do not show Beam leading its Chinese rivals. Best to treat this as an ambitious first entry, not a victory lap with the benchmark tables doing the driving.

What to watch

  • Whether Reflection releases Beam’s weights and technical report in October.
  • How independent evaluations compare its reasoning quality and compute use.
  • Whether the model’s performance holds up on real coding and agent workloads.

Discussion spark: When choosing an open model, would you take a lower-ranked system that claims to use much less reasoning compute, or wait for independent tests on your own workloads?

Sources and evidence

Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.