- Unsealed court filings quote a Microsoft exec calling AI scraping “the largest theft of labor in human history.”
- Microsoft privately described OpenAI‘s data practices as “theft,” the filings show.
- Both companies used paywalled New York Times content to build datasets.
- Internal warnings said the practice would gut publishers.
What Happened
Newly unsealed court filings show a Microsoft executive privately called AI scraping “the largest theft of labor in human history,” TechCrunch reported on September 17, 2026. The unredacted documents indicate Microsoft described OpenAI‘s data practices as “theft” even as both companies scraped paywalled New York Times content, built datasets from it, and warned internally that the practice would gut publishers.
Why It Matters
The filings surface in the copyright litigation between The New York Times and OpenAI and Microsoft, and they cut against the industry’s public fair-use defense. Internal acknowledgment that data practices amounted to “theft” — from the very companies mounting the fair-use argument — is the kind of evidence that can shape both the case and the broader legal question of whether AI training on copyrighted work requires a license.
Technical Details
Building a training dataset from paywalled content means circumventing access controls and copying material that publishers sell. The internal warning that this would “gut publishers” speaks to the market-harm factor courts weigh in fair-use analysis — the argument that AI systems trained on news can substitute for the news itself. The documents are contemporaneous internal assessments, which carry more evidentiary weight than after-the-fact positions.
Who’s Affected
OpenAI and Microsoft face damaging evidence in active litigation. The New York Times and other publishers gain leverage in their suits and licensing talks. Every AI developer relying on the fair-use defense is exposed if courts treat internal “theft” admissions as probative.
What’s Next
The filings will feed the ongoing case, where the fair-use question remains unresolved even as the US government has backed OpenAI’s position. How the court weighs these internal communications could influence the wave of parallel copyright suits against AI companies.