9/18/2026
Alex

Microsoft's Own Exec Called AI Scraping the 'Largest Theft of Labor in Human History.' The Story Is the Supply Chain Eating Itself

Here's the version of this story that got the headlines: a Microsoft executive privately called the AI industry's scraping of news content the "largest theft of labor in human history." That quote, buried in documents that were supposed to stay sealed, is the explosive part. But the documents that surfaced Thursday in the New York Times-led copyright suit against Microsoft and OpenAI tell a deeper story, and it is not about outrage. It is about a supply chain quietly eating itself.

The filings were unsealed as part of a motion for summary judgment in the multi-year case. Microsoft Director of Applied Science Brent Hecht appears repeatedly in the internal records, describing AI training on news as "an astonishing theft of unprecedented proportions" and, per the news plaintiffs, warning that the scale of it amounted to "the largest theft of labor in human history." When Microsoft's own applied-science lead writes that, internally, in a document that was meant never to see daylight, it means the people building these systems knew exactly what the cost looked like long before the rest of us did a thing about it.

OpenAI's own leadership said similar things in private messages. Nick Turley, the head of ChatGPT, wrote that publishers faced an "existential threat" from products that are "largely substitutive" and would "get more and more substitutive as they get better." Greg Brockman replied "ah nice" when informed that a staffer had found a way to get OpenAI's crawlers around the New York Times paywall. That is not the language of a company that believed it was exercising fair use. It is the language of a company that knew it was taking something.

The numbers are the part that should worry anyone running an enterprise that depends on published content. Microsoft's own data reportedly showed click-through rates for some news plaintiffs dropping 83 to 93 percent compared with traditional Bing search, and 51 to 94 percent for others. An internal Microsoft presentation from January 2024 described a "doom loop" that would "hurt the performance of our models and the entire web at the same time." The document is blunt about the stakes: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"

This is the tension that makes the case interesting rather than merely lurid. The fair-use defense has largely held so far in the courtroom; judges have been receptive to AI companies' arguments that training on copyrighted text is transformative. The Trump administration even chimed in this month with a brief backing OpenAI on that point. But the newly exposed internal admissions cut directly against the market-substitution branch of the fair-use test, which requires that the copying not harm the market for the original. When the companies' own slides show a 93 percent click-through collapse and their own executives write that the products are "substitutive, period," they hand the plaintiffs precisely the evidence a judge needs to find against them.

Under oath, Microsoft CEO Satya Nadella went further than the company's public line. He testified that anything paywalled "should be licensed by anyone who wants to use it for grounding or training," and said that had he known OpenAI had trained on paywalled content, he would have required OpenAI to retrain its models. The company's public spokesperson, meanwhile, dismissed Hecht's comments as "one employee's individual perspective" that "does not represent the company's views." Both things are now on the record, in the same case, and the gulf between them is the strategic problem Microsoft and OpenAI are trying to litigate their way out of.

There is a reason this matters outside the news business. Every enterprise that operates on data has now watched the economics of the thing it depends on get quietly cannibalized by the tools built on top of it. The "content supply chain" problem Microsoft's own document names is not unique to journalism. Any company whose products are trained on the work of others, or whose own work feeds someone else's model, is living inside the same loop. The plaintiffs put it in a single line: "The future not just of journalism but of responsible AI too depends on preserving incentives for humans to produce the creative works on which a healthy society depends."

The deeper lesson for the hardware and systems side of the industry is the same one playing out in every corner of the AI build-out. When a technology scales faster than the supply chain that feeds it, the short-term cheat rarely survives contact with the long-term bill. The teams that treat their data and content inputs as a real supply chain, with real licenses and real cost, are the ones that will not be scrambling when the courts, or the economics, finally catch up. The ones betting on "free" are betting on a liability they cannot see yet.

Building infrastructure on top of inputs you do not actually own is the kind of risk that does not show up on a spec sheet or in a dependency tree. At DMC, we help companies working at the edge of the AI and hardware build-out think through the full cost of what they are building before the bill arrives, whether that is a licensing problem, a supply chain constraint, or a compliance gap hiding in plain sight. If your roadmap has a dependency you are not sure you can defend, let's talk about stress-testing it before someone else finds the crack.