Skip to content
Benderson Media
Markets
AAPL $241.52 -0.38%
BTC $97,412 +3.21%
MSFT $478.90 +0.67%
ETH $4,128 +1.89%
GOOGL $182.34 -0.52%
TSLA $312.67 +4.23%
META $621.45 +1.05%
S&P 500 $6,142.80 +0.31%
NASDAQ $20,847.50 +0.78%
NVDA $183.06 +2.14%

News Publishers Sue OpenAI and Microsoft for Billions

News Publishers Sue OpenAI and Microsoft for Billions
Image: TechCrunch | Source

The Seattle Times and Newsday just joined a lawsuit wave that is not slowing down. Both publishers filed suit against OpenAI and Microsoft for scraping their journalism to train AI without a license or a check. The New York Times went first in December 2023, and according to court filings reviewed by Reuters, that case alone could expose OpenAI to billions in statutory damages. Every new plaintiff makes the fair use defense harder to sell.

What Is Actually Happening

Let’s get the basics straight. OpenAI and Microsoft built their AI products using billions of pages of text scraped from the internet. That text included copyrighted news articles from major publishers. Nobody asked for permission. Nobody paid for a license. The publishers found out, and they’re suing.

The New York Times fired the first major shot. According to the Times lawsuit itself, OpenAI’s models can reproduce Times articles nearly verbatim in some prompts, which goes well beyond what most legal experts consider fair use. Since then, the list has grown. The Denver Post, Chicago Tribune, and now Seattle Times and Newsday have all filed similar claims.

According to a 2025 Stanford Internet Observatory report on AI training data, news content represented a disproportionately large share of high quality text in the datasets used to train major language models. Publishers produce some of the most accurate, well edited, and densely informative content on the internet. That made them a prime target for scraping. According to Reuters, Microsoft is named as a defendant because it invested roughly 13 billion dollars into OpenAI and embedded that technology across its product line. If OpenAI infringed, Microsoft benefited directly.

The Money Play Nobody Is Talking About

Most people see copyright lawsuits and think “lawyers getting rich.” That’s the poor mindset reading of this situation.

Here’s my read. This is a battle over who owns the inputs to one of the most valuable industries ever built. AI companies trained on decades of human knowledge production and paid nothing for it. Now they’re worth hundreds of billions. The people who produced that knowledge are fighting to get paid retroactively. I think they have a real case.

OpenAI’s defense is built on fair use. That doctrine lets you use copyrighted material without permission in limited ways, mainly for commentary, education, or criticism. The problem is OpenAI isn’t commenting on news articles. It’s selling a product built on them. According to legal analysts at Bloomberg Law, courts weigh commercial benefit heavily in fair use analysis, and building a product you charge for is about as commercial as it gets.

According to PitchBook data, OpenAI’s valuation crossed 150 billion dollars in its most recent funding round. If publishers win even a fraction of what they’re seeking across all active suits, the payout would be enormous. More importantly, it would force every AI company to license training data going forward. That changes the economics of the entire industry permanently.

There’s a playbook here that smart operators should recognize. When a new technology company extracts value from an existing industry without paying, eventually the existing industry fights back. Music labels did it with Napster. Publishers did it with Google News. The settlement or court outcome shapes the market for years. We are at that inflection point right now for AI and media.

If you’re running a content operation of any size, your intellectual property portfolio is an asset on paper, not just a business function. Tracking legal and compliance spend cleanly matters when you’re coordinating across multiple contractors and law firms. Wallester’s business card platform makes it easy to separate and monitor those vendor categories, which matters as more content businesses start exploring their own IP protection options.

What This Means for You

If you create content for a living, this case matters even if you’ve never heard of Newsday. Here’s what I would do right now.

First, document your original content. Date stamps, author records, publication logs. If AI companies scraped your work, you’ll need proof you wrote it first. Most creators don’t have clean records. Fix that now.

Second, update your terms of service. Explicitly prohibit AI training use of your content. It may not stop scraping, but it weakens any fair use argument and strengthens your legal position if you ever need it.

Third, join your industry trade group. The News Media Alliance and similar organizations are the ones negotiating licensing deals on behalf of publishers. If OpenAI starts writing checks to settle suits, that money flows through those groups. Individual creators who aren’t affiliated with any organization will miss out.

Fourth, structure your content business like a real business. Clean books, separated accounts, documented payroll. Creators who operate informally are harder to include in class actions and have a harder time asserting IP rights. If you have contributors or contractors on your team, running payroll through Gusto keeps your records clean and your business credibly organized, which matters more than most people realize when legal disputes arise.

Fifth, start exploring proactive licensing. Several platforms are building registries where creators can license their content explicitly for AI training use at negotiated rates. That market is going to grow fast once court precedents set the price floor. Getting in early positions you to collect rather than chase.

The Bottom Line

OpenAI built a product worth hundreds of billions on content it didn’t pay for. That worked until the content owners noticed. Now they’re lining up in court, and they have real cases with real damages exposure. The free ride is ending. Either OpenAI starts licensing content at scale, or courts set the price for them. Either outcome means the AI companies that survived on free data are about to start paying for it. Plan your content business accordingly.

Frequently Asked Questions

Why are Seattle Times and Newsday suing OpenAI and Microsoft?

Both publishers allege that OpenAI and Microsoft used their copyrighted articles without permission or payment to train AI models. This follows the path set by the New York Times lawsuit filed in December 2023, which seeks billions in statutory copyright damages for similar claims.

What is the fair use argument and why is it weak here?

Fair use allows limited use of copyrighted material without permission, typically for education, commentary, or criticism. OpenAI has argued its training process qualifies. But according to legal analysts at Bloomberg Law, courts weigh commercial benefit heavily, and OpenAI sells a for profit product built on this content. That makes the fair use defense difficult to sustain.

Could these lawsuits actually cost OpenAI billions of dollars?

US copyright law allows statutory damages up to 150,000 dollars per willful infringement. With millions of articles potentially scraped without a license, the math grows fast. According to Reuters, the New York Times case alone carries this level of exposure. Multiple simultaneous suits compound the total liability significantly.

What should content creators do to protect their work?

Document your original content with clean date and authorship records. Add AI training prohibitions to your terms of service. Join trade groups that negotiate licensing deals on your behalf. Structure your business properly so you can assert legal rights if needed.

Will AI companies eventually pay for news content?

Some already are. Google struck licensing deals with several publishers before going to court. OpenAI has paid a handful of partners. If courts reject the fair use defense, licensing fees for training data become a standard cost of doing business in AI. That price gets passed to users and baked into product margins. The free data era is ending one lawsuit at a time.