Wild AI Web Text Is Starting to Poison Pretraining

Wild AI Web Text Is Starting to Poison Pretraining

4 min read

A Pangram-linked arXiv paper finds that AI-generated web text can help small, data-starved models briefly, then hurt human-text performance as it scales. The practical lesson is simple: separate human and AI corpora, measure them separately, and stop treating scraped web tokens as interchangeable.

TL;DR: AI-generated web text is not just “more data” for pretraining, because it can switch from useful to harmful depending on the model’s data budget and target distribution.

What happens when AI text gets scraped back into training data?

The primary source here is the arXiv paper “How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text,” which studies what happens when AI-written web pages quietly enter language model pretraining corpora.

The setup matters. This is not the usual synthetic-data experiment where a lab generates clean examples on purpose. It is “wild” AI text: content published on the open web, written for human readers, mixed across topics and formats, and scraped back into datasets without labels.

After applying FineWeb quality filtering, the paper reports that 27.5% of tokens from June 2026 web data were labeled as AI-generated by Pangram, rising to 31.1% by August. That is a detector-labeled estimate, not ground truth from every publisher on the internet. Still, the direction is hard to ignore. The web is becoming a feedback surface for models.

The team pretrained 800 language models while varying the ratio of added AI tokens to human tokens, then fit scaling laws against held-out losses on both human and AI-generated text. Their result is more interesting than “AI data bad.” For data-starved models, adding AI tokens initially lowers loss on human text. More data helps when you do not have much of it.

Then the curve bends. The benefit saturates, and eventually reverses into harm. For models already trained with high budgets of human text, AI tokens raise loss on human text almost immediately, while the same number of fresh human tokens keeps improving it.

clean human documents flowing into a model while mixed AI web documents create a loop that curls back and muddies the st

Why does the old scaling-law story break?

The paper says scaling laws like Hoffman et al. 2022, the Chinchilla-style compute and data tradeoff, do not predict this behavior. That makes sense. Classic scaling laws often treat tokens as if they mainly differ by quantity and distributional fit. This paper argues the token type itself can change value.

The proposed replacement has separate benefit and harm terms, which lets the value of an AI token change sign. That is the core idea. An AI token can be helpful early, neutral later, and harmful after that. Not because it is cursed. Because training on model outputs pushes the learner toward the statistical habits of prior models, not necessarily toward the human-text target you care about.

The paper reports that its new scaling law, fit on smaller models, predicted held-out human-text loss for models up to 3.6 times larger with 41% lower error than the best existing law across AI ratios. That is a meaningful claim, though still scoped to this experimental setup and detector-labeled web data.

The strongest operator takeaway is in the recommendations. The paper recommends filtering AI text when the target is human text, repeating human text before expanding a dataset with AI-generated web text, and reporting validation loss on human and AI text separately. That last part is underrated. If your validation set blends human and AI text, you can hide the exact failure you need to see.

Is AI-generated text still useful for training?

Yes, but the target matters.

The paper explicitly says AI text remains valuable when the target is AI text. That is not a loophole. It is a reminder that “quality” is always tied to what you want the model to do. If you are training a model to imitate assistant-style answers, summarize synthetic tickets, or operate inside an AI-heavy product environment, AI-generated examples may be relevant.

But if the goal is better human prose, better human reasoning traces, or cleaner grounding in human-authored web material, wild AI text is a contaminant once it crosses the helpful-data threshold. The danger is not dramatic model collapse overnight. It is slower: training runs that look bigger, cheaper, and more current, while quietly getting worse on the distribution users actually value.

Practitioner’s take: if you are building datasets, do not treat scraped web text as a single bucket anymore. Add AI-likelihood labels, keep human and AI validation sets separate, and run ablations before assuming more crawl data helps. The catch most teams miss is that filtering is not just a safety or copyright concern. It is now a model-quality control.