AI-Generated Books Are Growing Faster Than Reader Demand
Tech
AI
Publishing
Creator Economy
Copyright

AI-Generated Books Are Growing Faster Than Reader Demand

A 14,419-book study finds AI-heavy fiction gaining sales share while revenue per title falls. The causal story is less certain.

Uygar DuzgunUUygar Duzgun
Jul 28, 2026
Updated Jul 29, 2026
19 min read

To understand the market effect of AI-generated books, stop asking whether the average one is good.

The stronger question is whether supply is growing faster than reader attention and revenue. A July 2026 working paper reports exactly that pattern in a sample of 14,419 self-published genre-fiction ebooks: compared with 2023 Q1, the number of titles selling in a quarter grew 19.2 times by 2026 Q1, while unit sales grew 7.3 times and revenue 8.9 times. Books with substantial detected AI text sold worse on average, yet they gained sales and top-rank share as their numbers rose.

That is evidence of a supply shock. It is not proof that AI caused every decline, that every flagged book used AI, or that low-cost fiction leaves readers worse off. A second economics paper reaches an important counterpoint: more AI-containing books may add modest consumer value by serving niche demand, even while average usage per title falls.

The market can dilute author revenue and expand reader choice at the same time.

Reader level: Advanced. This article separates measured results, model-based estimates, legal interpretation, and practical decisions for authors, publishers, and platforms.

Contents

What the newest AI book study measured

Generative AI floods and dilutes the market for books, posted July 22 and revised July 26, studies self-published genre fiction released from January 2023 through March 2026. The researchers matched full book text to daily Amazon sales through June 29.

The analysis uses a proprietary panel maintained by a major publisher. The paper says the broader panel covers about 500,000 Amazon ebook identifiers representing at least 95% of daily ebook unit volume. Its study sample includes 14,419 titles across eight genre clusters after market-rank, genre, and text-availability filters.

The researchers split each book into chapters, removed front and back matter, and ran the text through Pangram 3.3. They placed titles into three bands:

No detected AI text: no analyzed window was flagged.
Light detected AI text: more than 0% and no more than 25% of windows were flagged.
Substantial detected AI text: more than 25% of windows were flagged.

These labels describe detector output. They do not reveal which model was used, whether the author disclosed it, how much editing occurred, or whether a human wrote the text. The paper tests alternate thresholds, but no detector turns authorship into an observed fact.

The study then compares sales, gross consumer revenue, rank positions, author output, Kindle Unlimited exposure, and distinctive phrase overlap. Its 90-day launch window makes older and newer release cohorts more comparable.

This design is unusually useful because it combines full text with transaction records. It is still observational. The safest reading is: the study measures a market pattern consistent with dilution and rules out several simple explanations; it does not identify a clean causal effect of AI adoption.

AI-generated books sell below their catalog share—and still matter

Books with substantial detected AI text were 20.0% of the sample but received 12.1% of launch-window unit sales and 11.3% of revenue. Books with no detected AI text were 62.9% of titles, 71.7% of sales, and 72.5% of revenue.

Detected-text bandShare of booksShare of salesShare of revenue
------:---:---:
No detected AI text62.9%71.7%72.5%
Light detected AI text17.1%16.2%16.2%
Substantial detected AI text20.0%12.1%11.3%

The average substantial-AI title underperformed. The category did not remain commercially irrelevant.

By 2026 Q1, titles with any detected AI text held 36.6% of observed sales and 38.5% of constructed Top-25 rank slots. The substantial-AI band alone held 17.5% of sales and 18.7% of those slots. The paper also documents a small set of high earners: its 15 highest-revenue substantial-AI titles each generated at least $220,000 in gross consumer revenue during the measured period.

Those figures do not establish author profit. Gross revenue precedes platform commissions, advertising, production costs, taxes, refunds, and other deductions. They do establish that “readers will ignore all of it” is not a defensible market assumption.

Average quality and aggregate impact are different variables. A low-cost producer does not need every release to succeed. It can publish many attempts, learn which covers and niches convert, and keep the small share that reaches scale. The relevant unit shifts from one book to a portfolio.

Supply grew faster than sales and revenue

The clearest result is the gap between title growth and market growth.

Indexed growth of released books, selling titles, revenue, and unit sales from 2023 Q1 to 2026 Q1
Indexed growth of released books, selling titles, revenue, and unit sales from 2023 Q1 to 2026 Q1

*Caption: Indexed change from 2023 Q1 to 2026 Q1 in the study sample. Released catalog is cumulative; selling titles, unit sales, and revenue are quarterly flows. Source: Chakrabarty et al. (2026).*

By 2026 Q1:

the cumulative released catalog was 38.35 times its 2023 Q1 level;
selling titles in a quarter were 19.21 times their earlier level;
quarterly revenue was 8.91 times its earlier level;
quarterly unit sales were 7.30 times their earlier level.

The cumulative catalog and quarterly flows are not identical measures, so the 38.35 figure should not be compared as if every released book competed equally in that quarter. The like-for-like quarterly comparison still shows the core imbalance: selling titles grew more than twice as fast as revenue and more than 2.6 times as fast as unit sales.

The paper then uses a fixed launch window to compare release cohorts. Mean launch-window revenue per selling title fell from $20,134 in 2023 to $17,312 in 2025, a 14% decline. For books with no detected AI text, it fell from $23,877 to $19,739, a 17.3% decline. Revenue per selling title fell in six of eight genre clusters across all books and in seven of eight clusters for the no-AI group.

This matters because it addresses a composition objection. The market average could fall simply because a flood of low-selling AI books entered the denominator. The no-AI cohort decline shows that the change reached books outside the substantial-AI band too.

Exposure patterns point in the same direction. The no-AI share of constructed Top-25 positions was 87.8% in low-exposure genre-months and 62.8% in high-exposure ones. In Kindle Unlimited-heavy genres, the no-AI lead over substantial-AI titles was 8.5 percentage points smaller for sales and 8.4 points smaller for revenue.

Amazon says Kindle Unlimited royalties come from a monthly global fund and depend on each title's share of normalized pages read. That creates an explicit shared pool. The study's stronger dilution pattern in KU-heavy genres fits the mechanism, but it does not prove that Kindle Unlimited caused the difference.

Why weak average quality does not prevent dilution

Creative markets allocate more than money. They allocate search visibility, recommendation slots, review attention, audience time, and editorial screening.

A new title can impose small discovery costs even when it sells nothing:

a search or recommendation system must rank it;
a reader must inspect or skip it;
competing titles move down or receive fewer impressions;
moderation and quality systems absorb more volume;
publishers and authors spend more to signal trust.

The effect accumulates across thousands of releases. This is why production scale matters independently of average product quality.

The study finds that 287 of 385 observed byline identities that continued publishing substantial-AI titles increased their monthly output after adoption. That 74.5% is conditional on authors who kept publishing in the category. It is not an adoption effect for all writers: only 311 of 824 identities in the broader event-time panel sat above the output-growth diagonal.

Recommended reading

The cost side helps explain the asymmetry. AI can compress drafting time, but it does not make generation free. More output still consumes inference energy, review time, asset work, platform capacity, and reader attention. The broader cost structure is covered in the resource cost behind additional AI output. A producer may rationally accept mediocre average performance when the cost of each additional attempt is low enough.

Recommended reading

Quality remains hard to define. Sales reward fit, packaging, timing, price, promotion, and existing audience as well as prose. A separate analysis of open and closed models shows why creative quality depends on the evaluation target. A market study can measure what readers buy; it cannot reduce literary value to revenue.

The reader-value counterargument

Market dilution does not imply that consumers receive no value.

The May 2026 revision of the NBER working paper AI and the Quantity and Quality of Creative Products examines a much broader Amazon ebook market. It uses a stratified sample of more than 330,000 releases representing about 10 million ebooks from 2020 through 2025, plus a 479,000-book census across eight subcategories.

Its main quality proxy is usage: the age-adjusted number of reader ratings, validated against estimated sales and supplemented with sales-rank and star-rating checks. That is a demand measure, not a judgment of literary merit.

The paper reports three findings that complicate a simple collapse narrative:

monthly ebook releases nearly tripled from 2022 to late 2025;
average usage per title fell, and books flagged as containing AI had lower usage on average;
the larger catalog increased the absolute number of moderately used books and did not displace release activity by authors active before the LLM period.

The authors calibrate a nested-logit demand model and estimate that AI books raised consumer surplus by about 7% in 2025. That estimate depends on the model's substitution structure and the assumption that observed usage captures consumer utility. It is not a direct survey result and does not measure author welfare, search cost, cultural value, or long-term market quality.

The two book-market papers can both be right. More niche products can improve the chance that a reader finds a specific trope or combination. The same expansion can lower revenue per title and make discovery more expensive for creators.

PerspectivePotential gainPotential loss
---------
ReaderMore combinations, faster supply, underserved nichesHigher search cost, uncertain provenance, repetitive catalogs
AuthorLower production cost, more experimentsLower revenue per title, harder discovery, imitation pressure
PlatformMore inventory and engagementModeration, ranking, disclosure, and trust costs
PublisherFaster testing and assisted workflowsBrand risk, rights uncertainty, weaker scarcity

The distribution decides who benefits. A positive estimate for aggregate consumer surplus does not compensate a particular author whose expected income falls.

What private AI fiction reveals—and what it does not

Published ebooks are only one route for AI fiction. AI Fiction in the Wild, revised June 23, studies 573,453 English-language conversations from the WildChat dataset, collected from April 2023 through May 2024 through free GPT-3.5 and GPT-4 interfaces.

The authors used a lexicon and an o4-mini classifier, then manually checked a random sample of 300 conversations. The fiction classifier reached 97% precision and 94% recall against the authors' consensus labels. It classified 195,271 conversations—34%—as involving fiction. Fanfiction appeared in 49% of the fiction subset, and a small group of power users produced much of the activity.

This supports a demand signal for immediate, customized, repetitive, and niche fiction inside a private chat. It does not measure published-book demand.

WildChat users opted into a public research interface without requiring an OpenAI account. The sample can overrepresent technical users, power users, boundary testing, and people seeking a free service. The researchers cannot observe whether generated stories were shared, purchased, revised into books, or read beyond the conversation.

The economic distinction is important:

Private generation collapses reader and producer into one person.
Marketplace publishing asks an unknown reader to choose the work among substitutes.

AI can satisfy the first use case without creating a sale in the second. That helps explain how AI fiction demand can be real while average revenue per published title declines.

The paper provides its classification materials and data references openly. Its companion repository is useful for auditing definitions and reproducing the analysis. Repository availability strengthens process transparency; it does not make the underlying WildChat population representative.

Copyright and detection complicate the market story

The July market study adds a provocative textual result. Among top-50 sellers, books with substantial detected AI text had 45.0% coverage by rare five-or-more-word expressions found in a small number of existing books but absent from a 4.7-trillion-token web snapshot. The corresponding no-AI group measured 37.7%.

Within the substantial-AI top-200, rare-expression coverage rose 7.6 percentage points for each tenfold increase in revenue. The slope was statistically significant at *p* = .001, but the interaction comparing that slope with the no-AI group was marginal at *p* = .063.

This does not prove that a particular passage was copied from a particular book. Aggregate phrase overlap can reflect training data, genre conventions, quotation, common source material, detector selection, or other pathways. The paper itself says the analysis cannot establish infringement.

A separate preregistered study, Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers, shows why the quality ceiling may change. Twenty-eight MFA-trained readers and 516 college-educated general readers completed 10,920 blind pairwise evaluations of human and AI excerpts written in the styles of 50 award-winning authors.

With ordinary in-context prompting, MFA readers strongly preferred human excerpts. After author-specific fine-tuning on complete oeuvres, the tested ChatGPT outputs reversed that preference on average for both writing quality and stylistic fidelity. Fine-tuned output also evaded the two tested AI detectors far more often.

The study covers excerpts of up to 450 words, not complete novels. Its reported generation costs exclude human steering, editing, long-form coherence work, and publishing. It tests imitation under unusually rich author-specific training, not ordinary consumer prompting. The result is evidence that detector accuracy and perceived quality can change with training and editing—not a forecast that AI novels will replace authors.

Recommended reading

That distinction matches a practical rule from the AI-music detection workflow: detector output is evidence, not ground truth.

Amazon KDP's current content guidelines require a publisher to inform Amazon about AI-generated text, images, or translations. They do not require disclosure for AI-assisted work, and the publisher remains responsible for intellectual-property and quality compliance. The public guideline describes disclosure to Amazon; it does not promise a reader-facing label.

The U.S. Copyright Office's generative AI training report treats market substitution, market dilution, licensing, and public benefit as separate considerations. It also says fair-use analysis remains fact-specific. A sales correlation, detector score, or phrase-overlap statistic cannot settle that legal test.

What the evidence does not prove

The current evidence supports a narrower conclusion than the headline of the July paper.

It does not prove causality

The book-market study observes exposure, sales, and cohort changes. Genre demand, pandemic-era behavior, advertising costs, pricing, subscription dynamics, platform ranking changes, and author entry can move at the same time. The exposure gradients and no-AI cohort decline strengthen the dilution interpretation, but no random assignment identifies the AI effect.

It does not observe author income

The revenue figures are gross consumer spending. Kindle Unlimited page reads are not separated from purchases in the proprietary panel. Pen names can split one producer into multiple bylines. Neither study gives a clean distribution of net author profit.

It does not validate disclosure

The researchers say none of the sampled books disclosed AI use in the data available to them. KDP's disclosure rule concerns information supplied to Amazon. A missing public label is not evidence that an author failed to notify the platform.

It does not turn detection into authorship

Pangram's reported validation results are not a labeled ground-truth test on this full book corpus. Edited, translated, formulaic, or model-specific writing may behave differently. Robustness across thresholds reduces sensitivity to one cutoff; it cannot eliminate systematic detector error.

It does not measure the whole book market

The July study focuses on self-published genre-fiction ebooks primarily sold through Amazon. Traditional publishing, nonfiction, print, audiobooks, libraries, direct sales, and other platforms may have different economics.

A decision framework for authors, publishers, and platforms

Do not reduce the decision to “use AI” or “ban AI.” Measure the variable that each party can control.

For authors: optimize reader retention, not release count

Track each release cohort for at least:

revenue per title after advertising and production costs;
sample-to-purchase conversion;
completion or normalized-page-read rate where available;
repeat readers and series continuation;
refund, low-rating, and quality-report rates;
editing hours per retained reader.

Set a stop rule before scaling. If output rises while retained readers, profit per editing hour, or series continuation falls, more releases are magnifying weak fit.

Use AI assistance where it preserves differentiated judgment: research organization, consistency checks, accessibility review, or controlled ideation. Keep provenance records for source material, model use, edits, and rights decisions. Those records will not resolve every legal question, but they make disclosure and correction possible.

Recommended reading

The operating principle is the same as using AI without erasing the author's own judgment: define the decisions the human must still own, then test whether the workflow preserves them.

For publishers: separate throughput from portfolio value

Evaluate AI-assisted projects against a human-authored baseline at the same genre, price, author stage, and marketing level. Report medians and distributions, not only total releases.

Require:

a rights and provenance record;
named human editorial responsibility;
long-form coherence and originality checks;
a market-overlap review before acquisition;
post-release cohort economics.

A faster manuscript pipeline has little value if acquisition, editing, positioning, and discovery remain the bottlenecks.

For platforms: test discovery quality under supply shocks

Volume limits and disclosure fields address only part of the problem. Ranking systems should be stress-tested as catalog supply grows faster than reader sessions.

Useful platform measures include:

impressions per selling title;
time to first meaningful engagement;
abandonment after preview;
duplicate or near-duplicate cluster size;
concentration of visibility across producer identities;
complaints and refunds by provenance band;
reader satisfaction after recommendation exposure.

Keep “unknown provenance” separate from “human-created.” Missing or stripped metadata is not evidence of human authorship. Preserve the source disclosure, transformation history, and confidence level through ingestion, moderation, and recommendation systems.

Conclusion

AI-generated books do not need to beat human books on average to reshape publishing.

The newest market evidence shows a supply shock: far more titles competing for sales, revenue, and rank positions that grew more slowly. The decline in revenue per selling title extends to books with no detected AI text, and high-exposure genres show the strongest losses. Those patterns are consistent with dilution.

They are not a causal verdict. Broader market data also suggests that a larger catalog can create modest reader value, especially for niche demand. Private fiction generation shows that some readers want combinations no publisher would efficiently supply.

The practical test is distributional: who gains from the extra variety, who pays the discovery and quality costs, and which metrics reveal the trade before the catalog scales?

Frequently asked questions

Are AI-generated books allowed on Amazon KDP?

Yes. Amazon KDP requires publishers to inform Amazon when a book contains AI-generated text, images, or translations. It does not require disclosure for AI-assisted work. Publishers remain responsible for intellectual-property and quality compliance.

Are AI-generated books profitable?

Some are. In the 14,419-book study, substantial-AI titles earned a smaller share of revenue than their catalog share, but a small group reached meaningful commercial scale. Gross revenue is not author profit, and the average result does not predict a specific book.

Can an AI detector prove that a book was AI-written?

No. It estimates whether text matches learned patterns. A detector can support an investigation or a market-level analysis, but it cannot establish authorship, disclosure, copying, or infringement on its own.

Claim checks

ClaimEvidence statusQualification
---------
Selling titles grew 19.21× while quarterly unit sales grew 7.30× and revenue 8.91×Supported by the July market paperIndexed within the study sample; observational; released catalog is cumulative
Substantial-AI titles were 20.0% of the sample but 11.3% of revenueSupported by the July market paper“Substantial AI” means more than 25% of text windows flagged by Pangram
Revenue per selling title fell 14% from the 2023 to 2025 launch cohortsSupported by the July market paperGross consumer revenue over a fixed 90-day launch window, not author earnings
AI books raised consumer surplus by about 7% in 2025Model-based estimate in the NBER paperDepends on a nested-logit calibration and usage as a utility proxy
Thirty-four percent of WildChat's English conversations involved fictionSupported by the WildChat paperOpt-in public interface sample; not representative of all readers or book purchases
Fine-tuned AI excerpts outperformed human excerpts in the tested preference studySupported in that experimental settingShort excerpts, author-specific complete-oeuvre fine-tuning, and blind preference—not full novels or sales
KDP requires disclosure of AI-generated but not AI-assisted contentSupported by current Amazon documentationDisclosure is to Amazon; the public page does not promise a reader-facing label
The evidence proves AI caused market dilutionNot supportedCurrent market designs are observational and cannot isolate every competing cause

Sources

Chakrabarty, T., Liu, X., Ginsburg, J. C., and Dhillon, P. — Generative AI floods and dilutes the market for books, arXiv, revised July 26, 2026.
Reimers, I., and Waldfogel, J. — AI and the Quantity and Quality of Creative Products: Have LLMs Boosted Creation of Valuable Books?, NBER Working Paper 34777, revised May 2026.
Gupta, N., Antoniak, M., and Walsh, M. — AI Fiction in the Wild, arXiv, revised June 23, 2026.
Chakrabarty, T., Ginsburg, J. C., and Dhillon, P. — Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers, arXiv, revised March 17, 2026.
Amazon Kindle Direct Publishing — Content Guidelines, checked July 28, 2026.
Amazon Kindle Direct Publishing — Royalties in Kindle Unlimited, checked July 28, 2026.
Gupta, N., Antoniak, M., and Walsh, M. — AI Fiction in the Wild companion repository, checked July 28, 2026.

Recommended for you

AI Reasoning Tokens: More Tokens, Better Results, More Power

AI Reasoning Tokens: More Tokens, Better Results, More Power

OpenAI and Anthropic make the tradeoff visible: better AI answers often need more reasoning tokens, more inference compute, and more electricity.

7 min read
GLM and Kimi vs Claude and GPT: Why Open Models Feel Constrained

GLM and Kimi vs Claude and GPT: Why Open Models Feel Constrained

My testing says GLM and Kimi follow known patterns while Claude and GPT explore. The benchmarks that measure this gap — and where it flips.

8 min read
How I Built an AI Content Pipeline That Writes Like Me

How I Built an AI Content Pipeline That Writes Like Me

I built an AI content pipeline that writes like me using author context, Search Console data, and real internal links.

9 min read