Skip to content
Benderson Media
Markets
AAPL $241.52 -0.38%
BTC $97,412 +3.21%
MSFT $478.90 +0.67%
ETH $4,128 +1.89%
GOOGL $182.34 -0.52%
TSLA $312.67 +4.23%
META $621.45 +1.05%
S&P 500 $6,142.80 +0.31%
NASDAQ $20,847.50 +0.78%
NVDA $183.06 +2.14%

Amazon Is Burning Books to Feed Its AI Models

Amazon Is Burning Books to Feed Its AI Models
Image: TechCrunch | Source

Amazon poured $4 billion into Anthropic according to Reuters, and now reports are surfacing that the company is acquiring rare physical texts and feeding them into AI training pipelines at scale. Some of those books will not exist after. The company that made its first dollar selling books is treating them as disposable fuel.

The Books Are Not Coming Back

Amazon started as an online bookstore in 1994. Jeff Bezos picked books specifically because they were easy to ship and hard to damage. Thirty-two years later, the company controls the largest e-book platform on the planet, a massive print book business, and is now acquiring physical texts to train AI models.

Reports from archivists and rare book dealers surfaced in 2026 showing that Amazon-affiliated buyers are purchasing collections of rare and out of print texts. According to Publishers Weekly, some of these collections include first editions, manuscripts, and texts with no surviving digital copies. After purchase, the books are scanned and some are reportedly destroyed rather than preserved or resold.

The Internet Archive has documented similar patterns across the book digitization industry. According to the Internet Archive’s 2025 annual report, fewer than 30% of digitized physical texts are returned to accessible archives after scanning. The rest move into private corporate collections or are discarded.

This isn’t an accident. It’s a strategy.

Data Is the New Oil and Amazon Is Drilling Everything

Here’s what most people miss about this story. They focus on the cultural loss. I get it. Losing a 200-year-old manuscript is genuinely bad. But if you only see the tragedy, you’re missing the bigger play happening right in front of you.

Amazon is building a moat. AI models are only as good as the data they train on. Public web data has been scraped to the bone. Wikipedia has been in every model since 2018. Reddit’s content quality has collapsed. The next competitive edge in AI training is proprietary text that no competitor can access.

Rare books are perfect for this. They contain knowledge, vocabulary, reasoning patterns, and historical context that doesn’t exist anywhere online. A 19th century engineering manual, a hand annotated legal text, a collection of out of print scientific journals. Amazon buys the physical copies, extracts the data, and its AI gets a training advantage that Google, Meta, and OpenAI can’t replicate.

According to a 2025 report from The Atlantic, top tier AI labs now spend between $100 million and $500 million per year acquiring proprietary training data through purchases, licensing deals, and direct digitization projects. Amazon sits at the high end of that range based on its disclosed AI budget commitments.

Meanwhile, rare book prices have surged. According to the Rare Books and Manuscripts Section of the American Library Association, average auction prices for 18th and 19th century American texts rose 34% between 2023 and 2025. Corporate buyers, not private collectors, are driving most of that demand.

The poor mindset says this is a scandal and signs a petition. The owner mindset asks a different question: who profits from controlling this data, and how do I get positioned on that side of the table?

Content creators are already feeling this shift. If you create video content explaining AI, finance, or tech news, tools like InVideo AI let you produce professional explainer videos fast without a production crew. The creators who can explain these stories clearly are capturing the audience that used to read long form journalism. That audience is growing fast and it needs human voices it can trust.

What This Means for You

If you create, curate, or sell knowledge for a living, this story matters directly.

Amazon is betting that proprietary training data will separate AI winners from losers in the next three years. If that bet pays off, the content you create today could become training data for models that compete with you tomorrow. That’s not a distant fear. Publishers, educators, researchers, and journalists are already watching AI generated content cut into their work.

Here’s what I would do right now.

First, understand the value of your own content. If you’ve written books, guides, courses, or long form material, that content has data value beyond its direct revenue. Some AI companies are actively paying for licensing deals. Know what you own and what rights you’ve signed away to publishers or platforms.

Second, build an audience that trusts you as a human. AI can replicate facts. It can’t replicate a specific person’s track record, relationships, and point of view built over years. Your reputation is the one asset that cannot be scraped and retrained.

Third, stay tooled up without overspending. If you want access to software that helps you move faster as an independent creator, AppSumo carries lifetime deals on content and productivity tools that let you compete without enterprise budgets. The creators who stay lean and well equipped will outlast the ones who get priced out.

The game is changing fast. The people who act like passive observers will wake up in three years wondering what happened to their audience, their income, and their industry.

The Bottom Line

Amazon went from selling books to destroying them. That’s not a metaphor. That’s a deliberate business decision. The company is betting that controlling rare, irreplaceable data creates an AI advantage nobody else can buy their way into. They’re probably right. The only question left is whether you’re watching this play out or figuring out where you fit in the new system before it solidifies around you.

Frequently Asked Questions

Is Amazon actually destroying rare books to train AI?

Reports from rare book dealers and archivists in 2026 indicate that Amazon-affiliated buyers are purchasing physical rare texts and digitizing them for AI training data. According to the Internet Archive, a significant portion of commercially digitized texts are not returned to public access after scanning. Whether physical copies are destroyed or retained in private storage after digitization varies by acquisition, but either outcome removes them from public access permanently.

Why does rare book data matter for AI training?

Rare books contain text that has never been digitized and is not available anywhere online. This gives AI models access to unique vocabulary, historical reasoning patterns, and specialized knowledge that competitors cannot access through standard data acquisition. According to The Atlantic, top AI labs are spending hundreds of millions of dollars annually to secure proprietary training data for exactly this kind of advantage.

Is Amazon destroying rare texts for AI training legal?

Purchasing and digitizing books you legally own is generally permitted, especially for texts in the public domain. Using that digitized content to train AI models is the subject of ongoing legal debate. Multiple lawsuits were filed against major AI companies between 2023 and 2025 over training data practices, and courts are still working through what counts as fair use versus infringement at scale.

What can authors and creators do to protect their work from AI training?

Register copyrights for any original work you haven’t already protected. Review contracts you’ve signed with publishers or platforms closely for language granting AI training rights. Organizations like the Authors Guild provide active guidance and are lobbying for stronger creator protections in AI training legislation at the federal level.

How does Amazon have an advantage over other AI companies in acquiring rare book data?

Amazon already controls the largest book retail and distribution network in the world, giving it access to physical inventory, seller relationships, and rare book market intelligence that pure tech companies don’t have. Its $4 billion investment in Anthropic according to Reuters signals that Amazon is serious about building AI capabilities that compound on top of assets its competitors simply can’t replicate quickly.