AI-generated text is becoming a larger share of the web, but feeding more of it into language-model training may eventually do more harm than good. In a study of 800 models, researchers found that synthetic text’s initial benefit for data-starved models can reverse as the budget for human-written text increases.
Watch Desk analysis
What happened
The researchers report that AI-generated text rose from 27.5 per cent of web corpora in June 2026 to 31.1 per cent in August. They trained 800 language models on different mixtures of AI-generated and human-written text, then proposed a scaling law that accounts separately for synthetic data’s benefits and harms.
The team also released WildAI, an 83-billion-token corpus, alongside code and models. The paper, “How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text”, was published on arXiv on 1 October.
Why it matters
The web is becoming a source of both training material and machine-made material. If synthetic text helps when human-written data is scarce but becomes damaging as training budgets grow, then “more data” is not a particularly useful plan without asking what that data is made of.
The findings concern the study’s training experiments and its proposed scaling law. They do not establish that every synthetic-data mix harms every model, or that the reported share applies to the whole web.
Our read
The useful contribution is the shape of the trade-off: synthetic text may be a stopgap, not an endlessly renewable substitute for human-written material. Researchers and model builders should be able to test that claim against WildAI and the released code, rather than treating “AI-generated” as one uniform kind of data. The internet’s content pipeline has acquired a feedback loop; the interesting question is how quickly it starts eating its own homework.
What to watch
- Whether other teams reproduce the reported benefit-to-harm shift.
- How the proposed scaling law performs across models and training setups.
- Whether WildAI’s data and code help researchers measure synthetic text in other web corpora.
Discussion spark: If synthetic text helps models when human-written data is scarce but can become harmful as training budgets grow, should model builders limit its use, or focus on better ways to identify and curate it?
Sources and evidence
- How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text (1 October 2026, 02:47 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.