Microsoft says virtually nobody was grabbing NYT articles through its chatbot — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Microsoft disclosed that fewer than 1% of 8.2 million Copilot chat logs—specifically chosen for keywords tied to plaintiffs' websites—reproduced at least 16 words from news content used to train the model.
- Microsoft said an expert for the Center for Investigative Reporting identified 51 instances of "substantial overlap" with CIR work in that dataset, while the authors' expert found only 24 responses with at least 30 matching words across all 8.2 million conversations.
- Microsoft claimed only 10 of 212 books evaluated showed any matches, framing these minimal overlaps as evidence that AI training on copyrighted material constitutes fair use because outputs serve "significantly different purposes" than the originals.
- Microsoft filed for summary judgment on Friday in the consolidated lawsuits brought by the New York Times, Center for Investigative Reporting, and Authors Guild against Microsoft and OpenAI, seeking to end the case before trial.
- The Trump administration filed a statement of interest this week in the New York Times case supporting OpenAI, adding a federal policy dimension to the legal battle.
Why it matters: If the judge grants Microsoft's summary judgment motion, it could short-circuit one of the highest-profile copyright fights over AI training data and set a de facto precedent that minimal regurgitation in chatbot outputs counts as fair use—tilting the playing field in favor of Microsoft and OpenAI against publishers seeking damages or licensing revenue.
Ask SkimNews



