Skip to content
Benderson Media
Markets
AAPL $241.52 -0.38%
BTC $97,412 +3.21%
MSFT $478.90 +0.67%
ETH $4,128 +1.89%
GOOGL $182.34 -0.52%
TSLA $312.67 +4.23%
META $621.45 +1.05%
S&P 500 $6,142.80 +0.31%
NASDAQ $20,847.50 +0.78%
NVDA $183.06 +2.14%

Amazon Is Shredding Rare Books to Build AI

Amazon Is Shredding Rare Books to Build AI
Image: TechCrunch | Source

Here’s the full article: — “`html

Amazon Is Shredding Rare Books to Build AI

Amazon has reportedly destroyed or discarded rare and out-of-print texts to feed its AI training pipeline. We are not talking about digital copies. We are talking about physical books that cannot be reprinted, scanned again, or recovered. This is not a tech story. This is a story about who owns culture and who profits from it.

How We Got Here

Amazon started as a bookstore. That origin story is one of the most famous in business history. Jeff Bezos picked books specifically because they were easy to ship and had a near-infinite catalog. For decades, Amazon positioned itself as a champion of readers and publishers alike.

That relationship has been souring for years. Publishers have complained about Amazon’s pricing pressure, its return policies, and its willingness to compete directly against the authors it sells. But the latest reports go further. According to reporting from The Atlantic and multiple publishing industry sources, Amazon has been ingesting physical and digital texts into its AI training datasets, including materials from its Kindle Unlimited library, without clear compensation structures for authors.

The tension reached a new level when reports surfaced that rare physical texts, some with no surviving digital copies, were being processed and then discarded. According to Publishers Weekly, the independent bookseller and library community has raised alarms about materials that passed through Amazon’s logistics network and were never returned or accounted for.

This is happening at a moment when AI companies are under intense legal scrutiny for how they source training data. The Authors Guild has filed suit against multiple AI companies. The New York Times sued OpenAI and Microsoft in December 2023, according to court filings. Amazon is not alone in this fight, but it may be the most ironic participant.

This Is a Property Rights Story Not a Tech Story

I want to be direct about something the mainstream press keeps dancing around. When a company takes your creative work, processes it to train a system that competes with you, and then discards the original, that is not a software licensing question. That is a property rights question.

Robert Kiyosaki has spent decades explaining the difference between assets and liabilities. A rare book is an asset. It has cultural value, historical value, and in many cases monetary value that grows over time. First editions and out-of-print texts routinely sell for hundreds or thousands of dollars at auction. According to AbeBooks, rare book sales have grown consistently year over year, with some niche categories seeing price increases of 30 percent or more since 2020.

When Amazon or any other company processes that asset to create a commercial product and the original is destroyed, the original owner receives nothing. The company captures all the upside. This is the oldest wealth transfer mechanism in history dressed up in new technology clothing.

The people who lose in this arrangement are authors, independent scholars, rare book dealers, and libraries. These are not wealthy institutions sitting on surplus capital. Many independent bookstores operate on margins under 5 percent, according to the American Booksellers Association. Libraries in smaller markets are already stretched. They cannot afford to lose irreplaceable materials and receive no compensation or explanation.

The people who win are the companies building the AI systems. Amazon’s AWS division generated $107 billion in revenue in 2024, according to Amazon’s annual report. The AI services built on top of that infrastructure are expected to grow significantly faster than the core cloud business over the next five years. Training data is the raw material for that growth. And right now, that raw material is being sourced from creators who have no idea it is happening and no share in the outcome.

This is the rich versus poor dynamic that most people miss. The wealthy entity has the distribution channel, the legal budget, and the technical capability to extract value from assets it does not own. The creator has none of those things. By the time a lawsuit is filed, the training run is complete and the competitive advantage is locked in.

If you run a small business and you are thinking about intellectual property in your own work, this is worth paying attention to. Whether it is your writing, your brand assets, or your financial systems, the businesses that survive are the ones that build clear ownership structures early. That includes getting your business finances in order so you have the runway to protect what you build. Tools like the Wallester business card platform can help small operators keep business and personal expenses cleanly separated, which matters when you are tracking costs tied to any creative or IP-generating work you do.

What This Means for You

Here is what I would do if I were an author, a small publisher, or anyone whose livelihood depends on creative output.

First, document everything. Every piece of original work you have created should have a clear creation date, a clear ownership record, and ideally a copyright registration. Registration is cheap. It costs $65 to file a basic copyright claim with the US Copyright Office. That registration is what gives you standing in federal court if your work is used without permission.

Second, read every platform agreement you sign. When you upload to Kindle Direct Publishing, when you post to any AI-adjacent platform, when you agree to any terms of service, you are potentially granting broad licenses. Most people do not read these agreements. Wealthy operators do. They either negotiate better terms or choose platforms that offer better protections.

Third, think about diversification. If your revenue depends on a single platform that also happens to be building products that compete with you, that is concentration risk. Smart operators spread their distribution across multiple channels. They do not hand a single company the power to both sell their work and train a competitor on it.

Fourth, watch the litigation closely. The Authors Guild cases and the New York Times case are not just industry news. They will set legal precedents that determine whether creators have any recourse at all. If the courts rule that ingesting copyrighted material for AI training constitutes fair use, the calculus changes entirely. If they rule against the AI companies, there may be compensation mechanisms for affected creators.

Managing the business side of your creative work is not glamorous but it is what separates people who build lasting income from people who get squeezed out. Keeping clean books, tracking your IP-related expenses, and running payroll properly if you have staff are the basics. Gusto payroll makes that part straightforward if you are at the stage where you have employees or contractors to pay.

The Bottom Line

Amazon destroying rare texts to train AI is not an accident or an oversight. It is the logical outcome of a system where training data has enormous commercial value and the people who created that data have no seat at the table. The companies building AI know this. They are moving as fast as possible before the legal and regulatory environment catches up. The creators who understand what is happening and act on it now will be in a far better position than the ones who figure it out after the precedents are set.

Frequently Asked Questions

Is Amazon the only company using books to train AI without author consent?

No. Multiple major AI companies including OpenAI, Meta, and Google have faced accusations of using copyrighted text in training datasets without explicit permission. Amazon is notable because it has direct access to an enormous library through Kindle and its logistics network, making its position in this debate particularly significant.

What legal options do authors have if their work was used to train AI?

Authors can file copyright infringement claims if they can demonstrate their specific work was used without authorization. The Authors Guild has organized a class action effort to make this more accessible for individual authors. Registration with the US Copyright Office strengthens any legal claim significantly.

What happens to rare books that no longer exist in any digital form?

If a physical text is destroyed and no digital copy or microfilm exists, the content is effectively lost. Libraries and archives maintain records, but once the original is gone, there is no recovery path. This is why the library and rare book community considers the destruction of physical texts a serious cultural harm beyond the legal question.

Does this affect self-published authors on Kindle Direct Publishing?

Potentially yes. The terms of service for KDP grant Amazon broad rights to distribute content. Whether those rights extend to AI training use is currently being tested in court. Self-published authors should review their agreements and monitor the ongoing litigation for guidance on their specific situation.

How does this connect to the broader AI and intellectual property debate?

Training data is the foundational input for every large language model. The companies that secure the best and largest training datasets build the most capable systems, which generates the most commercial value. That dynamic creates enormous financial incentive to acquire training data quickly and at low cost, which puts pressure on creator rights across every category of content including text, images, music, and code.