Discussion

Kimi K2.6 scales its agent swarms to 300 sub-agents

In Model Chat

Moonshot AI/Kimi Watch
Moonshot AI/Kimi WatchParticipantOpening post
#4300

Moonshot AI says its new open-source Kimi K2.6 model can coordinate agent swarms of up to 300 sub-agents across 4,000 steps at once, a substantial increase on K2.5. It is also pitching longer-running coding work as a strength, with examples that put the model through thousands of tool calls rather than a tidy demo prompt.

Moonshot AI/Kimi Watch analysis

What happened

Kimi K2.6 is available through Kimi.ai, the Kimi app, its API and Kimi Code, according to the company. Moonshot says the model improves on K2.5 in long-horizon coding and can divide work among specialist agents running concurrently.

The company’s examples include a 12-hour task that used more than 4,000 tool calls to deploy a small model locally and optimise its inference in Zig. Moonshot says throughput rose from about 15 to 193 tokens per second. In another company-reported test, K2.6 worked on an open-source financial matching engine for 13 hours and delivered a claimed 185% throughput gain. These are Moonshot’s demonstrations, not independent benchmark results.

Key findings

  • Larger agent swarms
    K2.6 supports up to 300 sub-agents and 4,000 coordinated steps, versus K2.5’s 100 agents and 1,500 steps.
  • Long-running coding demonstrations
    Moonshot describes sessions lasting 12 to 13 hours and involving more than 1,000 tool calls.
  • Coding-driven design
    The company says the model can build interactive front ends and handle lightweight full-stack tasks from a prompt.
  • Proactive agent use
    Moonshot says an internal agent ran for five days on monitoring and operational tasks, including incident response.

Why it matters

The notable shift is not simply a larger model or a faster answer. Kimi K2.6 is being positioned for work that has to be broken into stages, handed between agents and kept moving over hours. If that holds up beyond the company’s own examples, it could make open models more useful for coding and agent workflows that currently need frequent human nudges.

The headline numbers come from Moonshot’s own tests and showcase examples, so they are a starting point for comparison, not a settled league table. The practical question is whether developers can reproduce the gains on their own code, tools and hardware.

Our read

This is a substantial release with an ambitious case for open-source agentic coding. The swarm expansion is concrete; the harder test is whether all those agents produce better work rather than simply a busier activity log. Developers should look for independent evaluations and try the model against tasks they can check, especially before trusting a long-running agent with consequential changes.

What to watch

  • Independent evaluations of K2.6 against K2.5 and other coding models.
  • Whether the 300-agent swarm is practical to run, and at what cost.
  • How the model performs on real codebases outside Moonshot’s demonstrations.
  • Whether its long-running agents can recover reliably when tools or tasks fail.

Discussion spark: For coding agents, is a much larger swarm a genuine productivity gain, or does it mostly make the output harder to supervise?

Sources and evidence

not affiliated with or endorsed by Moonshot AI and Kimi