8.2 Million Conversations: Can Microsoft Sort Out Its AI Copyright Accounts?
In the series of copyright lawsuits filed by *The New York Times*, multiple news organizations and writers against OpenAI and Microsoft, Microsoft recently submitted a set of striking data to the court: among the approximately 8.2 million Copilot conversation records screened and identified as the most likely to contain the plaintiffs' works, only 24 responses had matches of more than 30 consecutive words with the books involved in the case; among the 212 books analyzed, only 10 had any matching content.
Microsoft intends to demonstrate that users do not use Copilot to read *The New York Times* for free or copy an entire book word for word. The occasional reproduction of a few lines of content identical to the original work by large language models should not overturn the judgment that the entire AI training process serves the purpose of transformative use.
The most critical part of this set of data is that AI copyright litigation is shifting from disputes over the principle of "whether the model has read the works" to an evidence war centered on "how much the model remembers, how much it outputs, under what conditions it outputs, and exactly who it replaces".
I. What the 8.2 million conversations "demonstrate"
The 8.2 million conversation records submitted by Microsoft are not ordinary samples randomly selected from all user conversations on Copilot. According to Microsoft's statement in the litigation, these records were specifically selected because they hit keywords related to the plaintiffs' news websites, so they belong to the set of conversations most likely to contain content from the plaintiffs' works.
In other words, this is not 8.2 million people randomly picked out of the vast crowd, but a group of the most suspicious people are first invited into the room, and then checked to see if they have the plaintiffs' articles in their hands.
Among this set of high-risk samples, the analysis results from the plaintiffs' experts are not limited to the single number of "24 times". Microsoft claims that 59,545 of the conversations have overlaps of at least 16 consecutive words with the news content used to support the model's responses, accounting for less than 1% of the 8.2 million records; experts hired by the Center for Investigative Reporting identified 51 responses that constituted "substantial overlap" with the organization's works.
The data on books is even lower. The writers' experts found 24 responses containing at least 30 matching words in the 8.2 million conversations, and among the 212 books tested, only 10 had any matching content.
These figures at least support Microsoft in putting forward a factual judgment. In real-world use, it is not a common phenomenon for Copilot to reproduce the involved news and books verbatim or in large paragraphs. The main purpose of ordinary users' conversations with the model is not to bypass paywalls, obtain full articles for free, or read a whole book continuously.
If the plaintiffs claim that Copilot is essentially a machine that distributes pirated articles and books to users for free at any time, the low overlap rate in 8.2 million conversations will obviously impact this description. This is exactly what Microsoft wants to prove.
II. The fact that the model rarely outputs original content does not mean there was no copying during training
Microsoft's argument is well versed in the "rules" of communication. With 8.2 million conversations and only 24 matches of more than 30 consecutive words from book texts, how can anyone claim that AI is massively replacing original works?
However, the issues in copyright law are far more complex, as AI may involve at least three relatively independent usage links.
The first link is that works are crawled, downloaded and incorporated into the training dataset; the second link is that the works or their features are used for model training, parameter adjustment and capability formation; the third link is whether the model outputs content identical or substantially similar to the original works after receiving user prompts.
The 8.2 million Copilot records mainly answer the third question, that is, what the model outputs in actual operation.
Some of the core claims put forward by plaintiffs such as *The New York Times* are at an earlier stage: whether Microsoft and OpenAI have copied their works without permission, incorporated them into training or retrieval-augmented systems, and used these works to form commercialized AI products.
Even if a machine reads the entire library and never recites any book in full to users, the fact that it "does not recite the original text" alone cannot directly lead to the conclusion that the initial act of obtaining, copying and using these books is necessarily legal.
Conversely, even if the model occasionally outputs 30, 50 or more words identical to the original work, the entire training process cannot be directly deemed infringing solely based on the number of words. Factual expressions, common phrases and headlines in news reports are not protected by copyright to the same extent as highly original plots and language in novels; the number of consecutive overlapping words is only a screening clue, and the type of work, the importance of the overlapping content, the user's prompting method and the overall impression of the output still need to be judged case by case.
Therefore, what Microsoft presented is not an answer sheet that can end the litigation, but a piece of evidence that is quite favorable to the plaintiffs' "output substitution" theory.
The low reproduction rate can weaken the allegation that the model is mass-producing pirated works, but it cannot automatically clear the doubts about the source of training data and the act of copying during training.
III. 16 words, 30 words and substantial similarity: whose measurement standard is more credible
There is another easily overlooked problem with this set of data: behind different numbers, the measurement standards applied are not the same.
For news content, the screening standard of "at least 16 overlapping words" is adopted; for books, the standard of "at least 30 matching words" is used; experts from the Center for Investigative Reporting further identified 51 cases of "substantial overlap" from the relevant records.
16 words and 30 words are quantitative standards, while "substantial overlap" needs to be judged in combination with the expressed content, work structure and protected elements. They can be used to identify suspected copying, but cannot replace each other, let alone simply equate all word overlaps with copyright infringement.
For example, an AI response may contain a dozen consecutive words identical to a news report due to citing the full name of an organization, event name and a set of factual data, but not all of these contents necessarily belong to original expressions; another response, although not copying 30 consecutive words, may have rewritten the most core and recognizable plot or viewpoint organization method of the work.
More importantly, how much original content the model outputs often depends on how the user asks questions.
The output results are obviously different when ordinary users ask for information, request news summaries, versus when users repeatedly ask the model to continue writing, recite content, ignore restrictions, or provide the full text of paid articles. Changes in the model version, system prompts, safety guardrails, retrieval sources and test time may all change the reproduction probability.
Therefore, the 8.2 million records are large in volume, but that does not mean they have no sample boundaries.
In the next step, the court will not only look at how many matches have been found, but also inquire about which time period these records come from, which version of Copilot they belong to, how the screening keywords are determined, what the deletion and deduplication rules are, why different thresholds are used for news and books, and whether the conversations not included in the analysis may produce different results.
In AI copyright litigation, the larger the sample size, the more important the methodology becomes.
A striking number can attract the public, but only a verifiable methodology can convince the court.
IV. How will the low reproduction rate affect the judgment of fair use
Microsoft hopes to use these data to support its fair use defense.
According to U.S. copyright law, fair use usually requires comprehensive consideration of the purpose and nature of the use, the nature of the original work, the quantity and substantiality of the portion used, and the impact on the actual and potential market of the original work.
The 8.2 million conversations are most likely to affect two aspects: the purpose of use and market impact.
If the vast majority of Copilot responses do not reproduce the original text, but use the model's capabilities to complete question answering, translation, summarization, writing, programming and information organization, Microsoft can accordingly claim that the common use of large language models is different from news publishing or book reading, and the works are trained to build a general-purpose tool capable of processing language and knowledge tasks, rather than reselling the original works through a new interface.
The low reproduction rate also helps Microsoft refute direct market substitution. If users cannot stably obtain full articles and books from Copilot, it is difficult to simply conclude that every AI conversation reduces a subscription, purchase or licensing transaction. The occasional output of a small amount of overlapping content by the model is functionally and commercially different from a pirated database that can provide the full text of original works on demand.
But the plaintiffs will also put forward another market theory: the harm caused by AI not only comes from verbatim copying, but also from the substitution of information demand and content entry.
Users do not necessarily need to read the full text. When Copilot can extract the core facts, exclusive information and main conclusions from reports and directly give a sufficiently usable answer, even without reproducing 30 consecutive words, it may cause users to stop clicking on news websites, stop viewing advertisements, or cease to be paid subscribers.
Therefore, what the court needs to judge in the future is not as simple as "whether AI has copied a complete article", but whether AI is essentially learning language and providing new tools, or using content formed by others' continuous investment to build a commercial system that gradually replaces the original content entry?
V. AI copyright cases are starting to shift from screenshots to logs
In past AI copyright litigations, plaintiffs often used specific prompts to make the model output content suspected of being identical to the original works, and then submitted several screenshots to the court; model companies would refute that these results were deliberately induced and cannot represent the normal operation state of the products.
This kind of debate easily leads to a situation where each side sticks to its own version of the story.
The emergence of 8.2 million real conversation records has pushed the debate from a small number of demonstrations to large-scale empirical analysis for the first time. The court can no longer only look at what the model "can" output in theory, but also analyze what it "usually" outputs in reality, how often a specific work is reproduced, which prompts are likely to trigger copying, and whether these outputs have actually changed users' content consumption behavior.
But this will also bring a new set of evidentiary challenges.
How long should AI companies retain user conversations? Should content that users have deleted still be preserved for litigation? How to desensitize millions of conversations? Can the plaintiffs obtain searchable structured data? Can model companies refuse to hand over the data on the grounds of trade secrets, privacy and processing costs? If the system is constantly updated, which version of the log is of evidential significance?
Prompts, model versions, invocation sources, retrieved content and generated results are jointly constituting the new electronic evidence in the AI era.
VI. AI copyright accounting cannot be calculated solely based on the number of copies
The data submitted by Microsoft this time may easily lead people to regard the AI copyright issue as a simple division problem: 24 divided by 8.2 million, the proportion is almost negligible, so is the copyright harm also negligible? This calculation method is overly simplistic.
If the dispute targets verbatim copying at the model output end, then the number of reproductions, the length of overlapping content, the scale of users and the degree of substitution for original work sales will of course affect the scope of infringement and the amount of compensation.
But if the dispute targets the unauthorized copying of works during the training phase, the unit of measurement may no longer be the number of outputs, but the number of works used, the technical and commercial benefits obtained, and the amount of training license fees that should have been paid.
If the dispute targets the retrieval-augmented function, the court also needs to distinguish whether the model "memorizes" the content from its parameters, or calls news websites, databases or other external content in real time when answering questions. Different technical paths correspond to different copying behaviors, licensing methods and responsible subjects, and cannot all be lumped into the general category of "whether AI training constitutes fair use".
Future AI copyright licensing will most likely not have a single uniform price.
Judgment from Zhichanli
Microsoft's 8.2 million conversation records are a favorable signal for AI companies, but they are far from the end of copyright litigation.
For right holders, litigation strategies also need to be adjusted accordingly. In the future, it is not enough to only prove that the model has "read" the works, or can occasionally generate several paragraphs of similar text. They need to further prove which specific link the copying occurs in, which outputs use protected expressions, how this use affects the subscription, advertising, licensing and content transaction markets, and establish a damage calculation model that can be reviewed by the court.
The next contest in AI copyright is not just about how the law interprets technology, but about whether data can prove the law.
This article is from the WeChat official account "Zhichanli" (ID: zhichanli), written by Shawn/MCP, and authorized for release by 36Kr.