September 18, 2026 — 21 Arrows
Microsoft Exec Called AI Training Scraping 'the Largest Theft of Labor in Human History'
Key takeaways
- Microsoft and OpenAI executives privately called AI training scraping 'the largest theft of labor in human history' while publicly defending the practice.
- Internal Microsoft documents warned of a 'doom loop' where AI scrapes content, trains on it, then replaces traffic to the original publishers.
- Both companies scraped paywalled New York Times content despite terms of service, according to newly unsealed court filings.
- The gap between private doubts and public messaging shows executives knew the consequences of their data practices.
- Business owners can block AI scrapers in their robots.txt file and should document original content to establish a paper trail.
The Development
On September 17, lawyers for the New York Times unsealed court documents (https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/) that had been redacted at Microsoft and OpenAI's request. The filings are part of the ongoing copyright lawsuit the Times filed against OpenAI years ago. Inside were private emails and internal documents from executives at both companies.
The language is striking. One Microsoft executive called the practice "the largest theft of labor in human history (https://www.404media.co/doom-loop-openai-and-microsoft-admits-llms-are-destroying-the-web-and-built-on-theft/)." An internal Microsoft document described it as "an astonishing theft of unprecedented proportions," noting that "millions of people around the world will soon consider large models 'hoovering up' all their work" without permission or payment.
Another Microsoft memo warned of a "doom loop." The company's AI products scrape content from websites, train models on that content, then generate answers that replace clicks to those same sites. The document stated: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"
According to the filings (https://www.404media.co/doom-loop-openai-and-microsoft-admits-llms-are-destroying-the-web-and-built-on-theft/), both companies scraped paywalled New York Times content, built training datasets from it, and warned internally that the practice would gut publishers.
Why It Matters to a Business Owner
If you run a business with a website, write content, or employ people who create anything that lives online, this matters because it exposes the economics beneath the AI boom.
The executives' private language contradicts the public position both companies have taken. In public, Microsoft and OpenAI have defended web scraping as fair use and a necessary practice for training AI. In private, they acknowledged it as theft and predicted it would destroy the businesses they were taking from.
That gap between internal doubts and external messaging is important. It means the companies knew what they were doing and what the consequences would be. It also means arguments about fair use and inevitable progress were made while executives privately used words like "theft" and "doom loop."
For business owners along the Grand Strand and beyond, the practical concern is this: if you publish content online, someone is likely training an AI model on it without your permission and without paying you. Then that model may answer questions using your expertise, reducing the traffic and revenue that would have come to you.
The unsealed filings show that Microsoft's own analysis confirmed this cycle. They scrape your work, they train on it, they serve answers based on it, and your traffic drops. The document called it a threat to "the entire web."
If you employ writers, photographers, designers, or anyone who creates digital work, their labor is part of what these executives privately called "the largest theft of labor in human history." That is not activist language. That is Microsoft's language.
What It Does NOT Mean
This does not mean AI is going away. Both companies have billions of dollars invested and will continue to build and deploy these tools. The lawsuit is ongoing, and no court has yet ruled on whether this practice constitutes copyright infringement or falls under fair use.
It also does not mean your content is safe just because you have a paywall or a terms-of-service page. The filings show (https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/) both companies scraped paywalled Times content anyway.
This is not a story about rogue engineers or an accident. These were considered business decisions made by executives who understood the implications. The documents show they knew, they did it anyway, and they asked the court to keep their private doubts sealed.
Finally, this does not mean you should stop publishing or creating. It means you should understand the landscape and make informed choices about how you protect and license your work.
A Practical Next Step
If you own a business website or publish content online, review your robots.txt file. This file tells automated crawlers which parts of your site they can and cannot access. It is not a legal barrier, but it is a documented statement of your intent.
Major AI companies have published the identifiers their scrapers use. OpenAI's is called GPTBot. Google has Google-Extended. You can block them in your robots.txt file. It will not undo past scraping, but it draws a line going forward.
If your business creates original content, photography, or design work, document it. Register copyrights where it makes sense. Keep records of publication dates. If this issue lands in front of a jury or a regulator, a paper trail will matter.
Finally, if you are evaluating AI tools for your business, ask the vendor where their training data came from. The unsealed documents make clear that some of the biggest players in the industry privately described their own data practices as theft. You should know what you are buying and what liability you may be assuming.
These filings will not be the last word. But they are the clearest window yet into what the people building these tools actually think about how they built them.
artificial intelligence · copyright · web scraping · business strategy · data ethics · content creation