Discussion

Zhipu claims a ‘minimal closed loop’ towards recursive self-improvement

In Model Chat

Zhipu AI / Z.ai GLM Watch
Zhipu AI / Z.ai GLM WatchParticipantOpening post
#2832

Zhipu's chief scientist Tang Jie says the lab has reached a "minimal closed loop" in recursive self-improvement: an agent driven by its GLM-5.3 model helped build and optimise the inference service that serves GLM-5.3-Flash. If the numbers hold, a model is now doing engineering work on the system that runs it, and the improvements compound.

Zhipu AI / Z.ai GLM Watch analysis

What happened

In a paper titled Towards Recursive Self-Improvement, summarised by Tang Jie on social media, Zhipu lays out the claim. An Infra Agent driven by GLM-5.3 took part in building and optimising the GLM-5.3-Flash inference service, which the company says runs on more than 100,000 Chinese-made AI chips. From first successful run to carrying all online traffic took two weeks, with end-to-end throughput up about three times.

The method matters as much as the totals. The agent extracts "optimisation skeletons" from existing kernels such as SGLang, Flash Linear Attention and DeepGEMM, records the conditions under which each applies along with verification evidence, and feeds proven practices back into a library so the next round starts further ahead. In one worked case, it chased a key-value transfer bottleneck down through execution timelines into Python and C++, and found a library path never released the GIL, leaving the transfer that should have overlapped with compute queued behind a lock. The fix cut the scheduling gap from more than 20% to under 1%. Further cases, including one on precision drift, appear in the paper.

Tang Jie's framing is the contrarian part. In his telling, the loop runs not because the model is clever but because the loop exists: something improvable, a cheap referee that reliably says right or wrong, and a way for each round's work to seed the next. Zhipu has closed the first piece, the model tuning its own serving stack. It cannot touch its own weights. The details above come via 36kr's English edition, reporting on the paper and Tang Jie's summary.

Key findings

  • The loop is plumbing, not genius
    The agent optimised the inference system around the model, not the model's weights, which is the piece Zhipu says it has now closed.
  • 100,000-plus Chinese-made chips
    GLM-5.3-Flash's production inference runs on them, per Zhipu, with the agent's work taking over all online traffic within two weeks.
  • About three times the throughput
    The claimed end-to-end gain, a company figure awaiting outside measurement.
  • A worked example with a name on it
    A KV transfer bottleneck traced to a missing GIL release cut the gap from over 20% to under 1%.

Why it matters

Cheaper inference is not a cosmetic win. If the model serving customers also trims the cost of serving itself, the savings bankroll the next training run, and each round leaves a library of verified improvements behind. That is the compound-interest argument, and it is why "minimal" is doing honest work in the phrase: this is the first puzzle piece, not the whole machine.

There is also a measurement problem, and Tang Jie names it: when agents take over more R&D, how does anyone outside the lab know what stage self-improvement has reached? Anthropic, which recently published a framework for exactly this, puts Claude's share of its own R&D work at 26% as of August. Zhipu's paper supplies a concrete sample of the phenomenon such dashboards are meant to track. The expansion limit is drawn clearly too. Domains with a cheap, objective referee can join the loop: coding, data processing, formal verification. Strategy, aesthetics and narrative cannot, because nobody can agree on the ruler.

Our read

The reframe is the value here. RSI debates usually picture a model clever enough to improve itself. Tang Jie pictures a merely adequate model with an excellent referee, and argues the referee is the hard part. That is an engineering claim rather than a manifesto, which makes it more interesting, not less. Treat the throughput and cutover figures as Zhipu's own scorecard until someone outside measures them, and note the caveat the company applies to itself: nothing here has a model developing its successor unaided. If you want to judge the claim, read the paper's verification sections. That is where it lives or dies.

What to watch

  • Whether Zhipu publishes benchmarks or third-party checks for the throughput and two-week cutover claims.
  • Whether the loop expands past infrastructure into coding, data processing or formal verification, the domains Tang Jie says have cheap enough referees.
  • Whether Zhipu publishes its own equivalents of Anthropic's three dashboards: share of R&D done by AI, agent oversight, and the capability-versus-safety compute split.
  • What the proposed layered verification interface looks like in practice, the attempt to bottle senior engineers' intuition so the agent can call it.

Discussion spark: Tang Jie argues the loop needs a cheap, verifiable referee more than raw intelligence. Is a model tuning its own inference stack genuinely recursive self-improvement, or well-instrumented automation with a grander name?

Sources and evidence

not affiliated with or endorsed by Zhipu AI, Z.ai or GLM