Microsoft Called AI Scraping 'Largest Theft of Labor' — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Brent Hecht, Microsoft's director of Applied Science, called AI scraping "the largest theft of labor in human history" in a January 2023 memo and later described Copilot's impact — a NYT click-through drop of up to 93% versus traditional Bing search — as a "doom loop" threatening the web.
- Nick Turley, OpenAI's head of ChatGPT, wrote internally that publishers face an "existential threat" from products "largely substitutive" for original reporting and "will get more and more substitutive as they get better."
- Satya Nadella testified that paywalled content "should be licensed by anyone who wants to use it" and that, had he known OpenAI scraped paywalled material, he would have "invoked [Microsoft's right to] require OpenAI to retrain its models."
- OpenAI's mid-training datasets contain more than 91,692 copies of works from the NYT, Daily News, and the Center for Investigative Reporting; a Common Crawl-derived dataset held more than 2 million nytimes.com documents, and a Project Mango dataset held at least 160,903 unique publisher works.
- Greg Brockman replied "ah nice" when told about a "hack to get around nytimes paywall," and OpenAI employees allegedly stripped copyright notices from training data to prevent models from reproducing them.
- The Trump administration filed a brief earlier this month defending OpenAI's unlicensed use of copyrighted material to train LLMs — even as these internal admissions contradicting that very position came to light.
Why it matters: OpenAI and Microsoft's fair-use defense hinges on training not substituting for or harming publishers' markets — but their own internal data showing a 93% NYT click-through drop and an "existential threat" memo directly undercut that position, giving the NYT and co-plaintiffs strong ammunition at trial.
Ask SkimNews



