Microsoft says virtually nobody was grabbing NYT articles through its chatbot

WorkAI.TV Editorial Desk
3 Min Read

Share with your CISO

Microsoft is pushing for summary judgment in its copyright fight with the New York Times, the Authors Guild, and the Center for Investigative Reporting, and it’s leaning on its own discovery data to do it. Out of 8.2 million Copilot chat logs selected specifically because they were likely to contain plaintiffs’ content, only 24 responses matched 30 or more words from books, and just 51 conversations showed substantial overlap with CIR journalism. Microsoft’s argument is that training on copyrighted material constitutes fair use when the resulting system serves a sufficiently different purpose from the original work.

What this means for your business

If your organization is currently deploying or evaluating AI systems that ingest proprietary or licensed content, the legal theory Microsoft is stress-testing matters more than its outcome. The fair use argument, which holds that transformative use of copyrighted material for training doesn’t require permission or payment, is the same legal foundation most commercial LLM vendors are standing on. A summary judgment win for Microsoft would give legal cover to that entire vendor class. A loss forces the industry into licensing negotiations that will almost certainly raise the cost of capable, current models.

The specific numbers Microsoft cites deserve scrutiny, not because they’re fabricated, but because they’re strategically framed. Eight-point-two million logs were chosen because they were most likely to surface infringement, and the overlap rates came back low. That’s either genuinely good news about how Copilot generates text, or it reflects how well Microsoft’s RLHF and output filters, the human-feedback training and safety guardrails baked into the system, suppress verbatim reproduction without eliminating the underlying reliance on training data. The NYT’s counsel isn’t wrong to note that the output question and the training question are legally distinct. Courts may agree.

The case that most CISOs should actually be watching isn’t the copyright exposure itself, it’s what a plaintiff victory would mean for data governance inside enterprise AI deployments. If courts rule that training on licensed or proprietary content creates ongoing liability, every organization running a fine-tuned internal model trained on customer data, contracts, or third-party research faces a version of the same question. I’d revisit this call if the judge denies summary judgment and the case proceeds to trial, because that’s when discovery into OpenAI’s training pipeline becomes public and the precedent risk sharpens considerably.

Based on reporting from Microsoft says virtually nobody was grabbing NYT articles through its chatbot, originally published 2026-09-04 12:05:00.

TAGGED:
Share This Article