LatestBC Conservatives down 16 seats, nine MLAs quit under Findlay

Canadian business, markets & economy · Saturday, 5 September 2026

Business

Microsoft says fewer than 1% of Copilot chats match news text; NYT disputes significance

In a September 2026 court filing, Microsoft revealed that only 24 of 8.2 million Copilot chat logs contain a 30‑word match to news content and fewer than 1% contain a 16‑word match, a figure the New York Times’ counsel says does not erase alleged copying.

Printed page of Microsoft's SEC filing showing Copilot phrase‑matching statistics

Fewer than 1% of the 8.2 million Copilot chat logs examined contain a 16‑word match to news content, and just 24 logs (0.03%) contain a 30‑word match, Microsoft disclosed in a September 2026 court filing. The figures, quoted by The Verge, are the first quantitative evidence offered by the software giant in its defence against a high‑profile copyright lawsuit brought by the New York Times and a coalition of book authors.

How the numbers were derived

As part of the discovery process, Microsoft supplied an expert hired by the publishing plaintiffs with the full set of 8.2 million Copilot chat logs generated in 2026. The expert counted how many logs contained phrases that matched published material at two thresholds: at least 16 consecutive words and at least 30 consecutive words. The filing reports that 59,545 logs met the 16‑word threshold – a proportion of less than 1% of the total – while only 24 logs met the 30‑word threshold, representing 0.03% of the sample.

Microsoft Copilot chat‑log match statistics (as disclosed in September 2026 filing)
Match length Number of logs Percentage of total (8.2 M)
≥16 words 59,545 <1%
≥30 words 24 0.03%
Source: The Verge (quoting Microsoft filing)

NYT counsel’s pushback

The New York Times’ lead counsel, Ian Crosby, responded to the filing in the same Verge article, arguing that the low percentages do not alter the core allegation that Microsoft and its AI partner OpenAI copied protected material. Crosby is quoted as saying the documents “show Microsoft and OpenAI stole from the Times regardless of the low match percentages.”

"The documents show Microsoft and OpenAI stole from the Times regardless of the low match percentages," Ian Crosby, lead counsel for the New York Times.

While the numbers illustrate the rarity of verbatim overlap, the legal question centres on whether any such overlap constitutes infringement, and whether the copying was systematic or incidental.

Why the figures matter for the AI‑publishing dispute

  • Evidence for defence strategy. Microsoft’s data aims to demonstrate that Copilot’s output rarely reproduces full sentences from news articles, suggesting that the system is more of a summariser than a copier.
  • Potential influence on settlement talks. Quantifying the extent of overlap could shape any future licensing agreement between AI developers and publishers, as parties negotiate compensation based on the volume of copied text.
  • Regulatory scrutiny. Lawmakers monitoring AI training practices may cite the 0.03% figure when debating whether existing copyright frameworks need to be tightened for large‑scale language models.
  • Impact on publishers. Even a handful of matches could involve high‑value news stories; the NYT’s argument is that any unauthorized use, however infrequent, undermines the economic model of news organisations.

Company background and broader context

Microsoft (NASDAQ: MSFT) is a U.S. technology giant headquartered in Redmond, Washington, with 221,000 employees and a market‑wide presence across software, cloud services and AI. The firm’s most recent 10‑K filing (filed 29 July 2026) reported net income of US$133.749 billion for fiscal year 2026 and total assets of US$758.376 billion. Its AI arm, in partnership with OpenAI, powers the Copilot assistant that is at the centre of the current litigation.

OpenAI, led by CEO Sam Altman and based in San Francisco, supplies the underlying large‑language model that Microsoft integrates into Copilot. The dispute pits the two companies against traditional news publishers who claim that training data and generated outputs infringe on copyrighted articles.

Timeline of disclosure

  • 4 September 2026: The Verge publishes an article quoting Microsoft’s court filing, revealing the < 1% and 0.03% match rates.
  • Earlier 2026: The New York Times files a lawsuit alleging that Copilot copies its articles without permission.
  • 2 September 2026: Microsoft files an 8‑K with the SEC, though the filing does not contain the AI‑related numbers; the court filing is separate.

What remains unknown

The filing does not break down which specific articles or topics appear in the 24 logs that meet the 30‑word threshold, nor does it disclose the commercial value of those matches. Neither Microsoft nor the NYT has provided a projection of how the numbers might change if a larger sample of logs were examined, or how the courts will weigh statistical rarity against the legal standard for infringement.

Outlook for the sector

Analysts watching the AI‑copyright frontier note that the disclosed figures are likely to become a reference point in future disputes. If courts accept the argument that sub‑1% overlap is de minimis, AI developers may face fewer licensing demands. Conversely, if the NYT’s position that any unauthorized copying is actionable gains traction, publishers could push for broader data‑use restrictions, potentially reshaping how large language models are trained.

For investors, the immediate market impact is muted; the numbers are technical and do not alter Microsoft’s earnings outlook, which remains anchored in its $36.148 billion 2026 revenue (as reported in the 10‑Q for fiscal year 2011, included for historical context). However, the litigation risk premium could be reassessed as the case proceeds, especially if the court rules that even minimal verbatim matches constitute infringement.

What comes next

The next procedural step is a court hearing on the admissibility of the match‑rate evidence. Both sides are expected to present expert testimony: Microsoft will argue that the low incidence demonstrates good‑faith compliance with copyright law, while the NYT will likely focus on the qualitative significance of the few matches that do exist. The outcome will inform not only the settlement prospects for this case but also the broader policy debate on AI training data and copyright protection.

Until a judicial determination is made, the numbers remain a focal point for both legal strategy and public discussion about the balance between innovation and intellectual‑property rights.

About the author

Lucas Bennett

Reporting for CityAM Canada on business and the wider Canadian economy.

All work by Lucas Bennett ›