Huawei says its OceanStor M900 can give AI inference clusters a shared pool of context memory at petabyte scale, with terabyte-per-second performance. The pitch is aimed at a growing infrastructure problem: long-context AI systems need somewhere to keep the KV cache when GPU memory is too small and expensive.
Watch Desk analysis
What happened
Huawei announced the M900 at HUAWEI CONNECT 2026, according to AI Magazine. The system is designed to pool storage across a cluster, using Huawei’s UnifiedBus interconnect and a storage architecture the company says connects NPUs to SSDs in one hop.
Huawei claims the design can increase available KV cache per NPU from gigabytes to terabytes. It also says the system cuts access latency by 90%, doubles inference-cluster token throughput and halves time to first token in typical AI programming scenarios. Those are vendor-reported performance claims, not independently established results.
Why it matters
KV cache holds information a model needs during inference. As context windows and agent conversations grow, keeping more of that working memory close to the processors can affect how many requests a cluster handles and how quickly it responds. Huawei is making the case that storage should be part of AI infrastructure planning, rather than a distant cupboard for data nobody needs until later.
The M900 is aimed at hyperscale data centres, so this is not a plug-in upgrade for a local workstation. But the underlying trade-off matters well beyond one product: more capacity can be useful only if moving data in and out does not become the new bottleneck.
Our read
This is a consequential infrastructure bet, and the idea is more interesting than another promise to throw GPUs at the problem. The headline gains are Huawei’s claims; buyers will need comparable workload results, pricing and deployment details before treating them as a shopping list.
What to watch
- Independent measurements of latency, throughput and time to first token on real workloads.
- How the M900’s capacity, interconnect and cost compare with other ways of serving KV cache.
- Whether the claimed SSD endurance gains hold under production write patterns.
Discussion spark: If long-context AI is running short of working memory, should data-centre operators invest in shared storage layers like Huawei’s M900, or keep prioritising faster, larger GPU memory?
Sources and evidence
- How is Huawei Addressing the Long Context Memory Problem? – AI Magazine (9 October 2026, 07:39 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.