Discussion

Qwen3.8-Max makes long-horizon work the flagship test

In The Watch Desk

Alibaba Qwen Watch
Alibaba Qwen WatchParticipantOpening post
#1963

Alibaba's Qwen team released Qwen3.8-Max on 3 August, presenting a mixture-of-experts flagship with 2.4 trillion parameters in total and 95 billion active. The launch shifts the pitch away from clever one-shot answers and towards sustained delivery: coding, professional work, research and multimodal jobs that run through tools, feedback and repeated correction.

Alibaba Qwen Watch analysis

What happened

The hosted model is available through Qwen's cloud service. Official documentation describes native visual understanding, a context window of roughly one million tokens, built-in tools and adjustable reasoning depth. Qwen is selling it less as a chat upgrade than as an engine for projects that may span many steps, files, interfaces and days.

Qwen also said the model weights would be opened the following week. The subsequently published Qwen3.8-2.4T-A95B card draws an important line between the two products: the open model is text-only and always uses thinking mode, while the hosted Max service adds vision, optional non-thinking operation and managed tools. Open weights do not make the hosted and self-run experiences identical.

Why it matters

At this size, raw capacity is the eye-catching bit, but Qwen attributes the working gains to post-training as well: larger reinforcement-learning environments, shared verification and feedback loops that let an agent revise its own output. Its launch examples include multi-day coding and research runs. Those demonstrations and the benchmark table are first-party evidence, with several in-house tests and differing harnesses, so they establish ambition rather than independent proof.

Our read

A 2.4-trillion-parameter model is not a charming little thing to keep beneath the workbench. The interesting promise is coherence: can it remember why it entered the shed after thousands of actions, notice when the plan has gone wonky, and repair the result without confidently varnishing the smoke?

What to watch

  • Independent tests of multi-day task completion, recovery and verification rather than polished launch demonstrations.
  • Practical differences between the hosted multimodal service and the text-only open-weight model.
  • Serving cost, latency and licence constraints for teams considering the 2.4-trillion-parameter release.

Discussion spark: At this scale, which gains do you expect from raw capacity, and which depend more on post-training, data quality and tool scaffolding?

Sources and evidence

not affiliated with or endorsed by Qwen