Microsoft: Copilot Rarely Copies News, Logs Show — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Microsoft submitted filings arguing Copilot rarely reproduces copyrighted material, backed by 8.2 million chat logs shared during discovery with plaintiffs' experts.
- The analysis found fewer than 1% of chat logs regurgitated at least 16 words matching news content, with 59,545 of 8.2 million conversations showing such overlap.
- Center for Investigative Reporting's expert identified 51 instances of "substantial overlap" with CIR work within Microsoft's dataset, a figure Microsoft itself disclosed in its filing.
- The authors' expert found only 24 responses across 8.2 million Copilot conversations containing at least 30 matching words, and just 10 of 212 evaluated books showed any matches.
- Microsoft is urging the judge to issue summary judgment, arguing that even rare reproduction does not undermine the transformative purpose of LLM training and should count as fair use.
- The New York Times rejected Microsoft's framing, with lead counsel Ian Crosby saying discovery proves Microsoft and OpenAI "stole" from the newspaper to build competing products.
- The Trump administration filed a statement of interest this week supporting OpenAI in the NYT case, adding a government weight to the fair-use debate.
Why it matters: Microsoft is using empirical chat-log data to press for an early dismissal in a consolidated copyright fight that could define whether AI training on journalism qualifies as fair use. If the judge grants summary judgment, publishers and authors lose their leverage to extract licensing deals from Microsoft and OpenAI; if not, the case proceeds toward trial with discovery exposing more internal details about how the models ingest news content.
Ask SkimNews



