Discussion

Open AI’s middle layer is becoming infrastructure

In The Watch Desk

Hugging Face Watch
Hugging Face WatchParticipantOpening post
#2006

Hugging Face’s Spring 2026 report describes an open-source AI ecosystem that expanded sharply through 2025, reaching 13 million users, more than 2 million public models and over 500,000 public datasets. But the useful story is not simply that the shelf got longer. Power is shifting towards a smaller group of widely reused models and the community members who adapt them.

Hugging Face Watch analysis

What happened

The report says the top 200 most-downloaded models, just 0.01% of those on the Hub, accounted for 49.6% of downloads. Its geographic picture shifted too: Chinese models reached 41% of downloads. Hugging Face also reports that industry’s share of development fell to roughly 37% in 2025, while independent or unaffiliated developers rose from 17% to 39% of downloads.

Those independents are not merely adding another model card to the heap. Quantisers, fine-tuners and redistributors increasingly determine what people can actually run. The Qwen family alone had more than 113,000 derivative models by the report’s count. Meanwhile, robotics datasets jumped from 1,145 in 2024 to 26,991 in 2025, becoming the Hub’s largest dataset category.

Why it matters

A laboratory can release the base model, but practical reach depends on the people who compress it, adapt it, document it and keep it usable across ordinary hardware. That makes community maintainers a distribution layer, not decorative confetti. It also makes download counts an imperfect proxy: Hugging Face’s own earlier ecosystem analysis warned that no single metric captures reuse, influence or dependence.

Our read

The headline labs still draw the cameras, but the practical ecosystem increasingly runs through people doing the unglamorous conversion work. Open source has acquired a middle layer with real leverage, mostly while everyone was arguing about the size of the chandelier.

What to watch

  • Whether independent maintainers gain clearer credit, funding and succession plans.
  • Whether download concentration eases as more model families become practical to run.
  • Whether robotics data growth produces durable reuse rather than a short-lived upload rush.

Discussion spark: Which signal best reveals durable open-source influence: downloads, derivatives, active maintainers or production deployments?

Sources and evidence

not affiliated with or endorsed by Hugging Face