AI Books Win Amazon by Volume, Not Quality: What the Slop Data Shows
A new full-text analysis of 14,419 self-published Amazon books finds AI-heavy titles taking top ranks and diluting revenue per book, reshaping a creative market through scale rather than quality. Here is what the numbers mean for anyone selling or writing.
TL;DR: AI-written books are not winning Amazon because they are good, they are winning because there are so many of them that they crowd the shelf and shrink the money everyone else earns per book.
The paper is “Generative AI floods and dilutes the market for books,” posted to arXiv under both cs.AI and cs.CL. It runs full-text AI detection across 14,419 self-published genre-fiction titles sold on Amazon from 2023 to 2026, matched against daily sales records through June 2026. None of those books disclose AI involvement. That last detail matters: this is not a survey of people who admit to using ChatGPT. It is a detection study on a catalog where nobody had to tell the truth.
The headline most people will take away is “AI slop is beating real authors.” That is not quite what the data says, and the gap between the two readings is the whole story.
What did the study actually find?
Two facts sit at the center. First, books with substantial detected AI text (the authors set the bar at more than 25% of the text) make up a large share of the catalog but a smaller share of sales. So on a per-book basis, AI-heavy titles still underperform. The slop-skeptics were partly right: buyers do not love this stuff.
Second, and this is where it gets uncomfortable, those same AI-heavy books are winning a growing share of total sales over time and grabbing more of the scarce top-rank positions that used to belong to books with no detected AI text. Underperforming per title, gaining overall. The only way both are true at once is volume. There are simply a lot more of them.
The market-level number makes it concrete. Over the study period the count of books with observed sales in a quarter grew 19.2-fold, while quarterly revenue grew only 8.9-fold. Read that twice. The number of selling books more than doubled relative to the money. Revenue per selling book fell across most genres.

That is dilution, plainly. Same pie, more forks. And the forks that showed up fastest were the cheap ones.
Why does Kindle Unlimited make it worse?
The pattern is not uniform. Books with no AI text lose the most ground in genres where AI has spread the furthest, and most of all where Kindle Unlimited availability is high.
That mechanism is worth sitting with, because Kindle Unlimited changes the economics of reading. Under a subscription-and-page-reads model, a reader’s cost of trying another book is basically zero. You are not deciding whether a title is worth ten dollars. You are deciding whether it is worth the next twenty minutes, and if it is not, you swipe to the next one at no extra charge. In that environment, quantity is a strategy. Flood the pool with enough passable titles and some of them get read regardless of whether any single one would survive a paid purchase decision.
So the venue where trying a book is free is exactly the venue where sheer output pays off. AI lowers the cost of producing a book to near zero on the supply side. Kindle Unlimited lowers the cost of sampling one to near zero on the demand side. Put those together and you get a market that rewards throughput over craft. Not because readers got dumber. Because the friction that used to filter output disappeared on both ends.
Are these books just copying existing writing?
Here is the finding I keep turning over. Among top-selling books, the ones with substantial AI text draw on more distinctive language from existing books than the no-AI titles do. And for those AI-heavy top sellers, that overlap rises with revenue. More borrowed distinctive language, more money. The authors do not detect that gradient for books with no AI text.
Put bluntly: within the AI winners, the ones that sound more like existing published work tend to sell more. That is what you would expect from a system trained to reproduce patterns from a corpus, then pointed at a genre with strong conventions. Romance readers want the beats. LitRPG readers want the mechanics. A model is very good at giving you the beats, and the closer it hews to what already worked, the better it does.

This is not a smoking gun of infringement, and the paper does not claim it is. But it connects directly to the legal question the authors flag: the market-effect prong of the fair use defense. One of the tests for whether training or output infringes is whether it harms the market for the original work. If AI titles that lean harder on existing books’ distinctive language earn more, and if the overall market is diluting revenue for human authors most in the genres where AI is thickest, that is exactly the kind of concrete market harm plaintiffs have struggled to document. A dataset of 14,419 books with matched daily sales is a different animal than vibes and anecdotes in a courtroom.
I want to be careful here. Full-text AI detection is imperfect, and any 25% threshold is a judgment call. A human who edits AI drafts heavily might read as human; a plain human writer with a formulaic style might trip a detector. The paper’s strength is scale and the sales match, not certainty about any single book. Treat the individual labels as noisy and the aggregate pattern as the real signal.
What does this mean if you sell books?
The optimistic read, and I do think there is one, is that quality still wins per book. AI-heavy titles underperform on a per-title basis. The problem is not that the machine writes better than you. The problem is that it writes more than you, cheaply, in a store that no longer charges readers to try the next thing.

So the operator move is not “write faster to compete on volume.” You will lose that race to something with no marginal cost. The move is everything that does not dilute: a name readers seek out on purpose, a mailing list that bypasses the algorithm’s flooded shelf, a series with continuity that a batch generator cannot fake, a voice distinct enough that sounding like existing books is not the win. The study’s own finding cuts your way here. The AI winners succeed by resembling the corpus. Resembling nothing else in it is the one position a flood cannot occupy.
The catch most readers will miss: this is not really a story about books. It is the first clean, large-scale measurement of what happens to any creative market when production cost collapses and the distribution channel does not filter for effort. Stock images, music, short video, template code, marketing copy are all sitting in the same current. Watch the revenue-per-unit line, not the unit count. The unit count always goes up. Whether the money follows is the number that tells you what kind of market you are actually in.
Primary source: “Generative AI floods and dilutes the market for books,” arXiv cs.AI / cs.CL.