How will things pan out when large language models no longer redirect users back to content platforms?
In the past, search engines also crawled articles, answers, and web information on content platforms.
However, search engines typically only provide titles, summaries, and links. When users want to access the full content, they still need to click the link and navigate to the original platform where the content was generated.
Search engines crawled the content, but they also sent users back to the original platforms.
Now, large language models (LLMs) also obtain content from the internet, yet they may complete retrieval, summarization, rewriting, and response generation entirely within their own interfaces. After a user poses a question, they receive a complete answer directly, eliminating the need to visit the original content platform at all.
Shifting from "helping users find content" to "directly replacing platforms to provide content" changes not just the technical form, but the distribution of traffic, content, and commercial benefits.
Therefore, the real conflict between knowledge platforms and LLMs may not simply be whether LLMs have crawled content, but rather, after an LLM uses a platform's content to answer a user, why does it no longer send the user back to that platform?
(This article discusses a potential industry dispute and does not refer to any specific case or enterprise.)
1. Both search engines and LLMs crawl content, but their outcomes are completely different
From a technical perspective, both search engines and LLMs can access publicly available content on the internet.
Yet their positions within the content industry are not the same.
Search engines build indexes. They tell users where the required information might be located, and then direct users to the original websites via links. As a result, content platforms gain traffic, ad impressions, membership conversions, and transaction opportunities.
Under this model, while there are occasional conflicts between search engines and content platforms, a largely interdependent relationship has been formed —
Platforms provide content, and search engines provide traffic; search engines use platform content to build indexes, and then send users back to the platforms.
LLMs have broken this cycle.
In the past, users needed to ask a question — view search results — click a link — enter the platform — read the content. Now, this process may be reduced to: ask a question — get an answer — finish.
LLMs no longer simply tell users "where the answer is"; they deliver the answer directly to the user.
Reading, comparing, evaluating, and summarizing tasks that once had to be completed on content platforms can now all be done within the LLM's interface. Users gain convenience, but the content platform disappears from the entire service chain.
This is what knowledge platforms are truly concerned about:
What LLMs take away may not just be a single article, but the connection between the platform and its users.
2. What platforms lose is the entire content cycle
For content platforms, traffic is not an isolated number.
After users enter a platform, they read, comment, like, bookmark, and follow authors, and may also purchase memberships or use other services. All these behaviors together constitute the platform's commercial value.
The platform then uses the revenue generated to maintain its technical systems, moderate content, optimize recommendations, incentivize creators, and promote the continuous production of new content.
This originally forms a cycle: Creators produce content — the platform organizes content — users visit the platform — the platform gains revenue — creators continue producing.
When LLMs extensively use platform content but keep users within their own products, this cycle can be severed.
LLMs gain the ability to answer questions, users get efficient answers, but content platforms may lose:
User visits and dwell time;
Opportunities for ad impressions and membership conversions;
Community interactions formed by comments, likes, and follows;
Continuous insights into user demands and content feedback;
Opportunities for creators to gain attention and revenue;
The commercial foundation to continue organizing and producing high-quality content.
Therefore, the harm claimed by content platforms is not just "one of my articles was used", but rather:
LLMs use the content accumulated by platforms to build new knowledge services, while gradually depriving the original platforms of the ability to sustain their content ecosystems.
If this model continues to develop, a paradox may eventually emerge. LLMs become increasingly dependent on high-quality content, yet the platforms and creators who produce that high-quality content find it harder and harder to get rewarded.
3. Public accessibility does not mean unrestricted utilization
Faced with doubts from content platforms, LLM companies may argue that this content was originally published publicly on the internet, and ordinary users can read it — so why can't machines access it?
This question cannot be simply answered with a "yes" or "no".
Public accessibility means that content can be browsed by users within the scope of the website's normal settings, but it does not inherently mean that any commercial entity can obtain it in bulk, store it long-term, and use it to build competitive services without restrictions.
To judge whether a data acquisition behavior is legitimate, at least the following factors need to be considered:
Whether the content obtained is ordinary public content, or content that requires login, is restricted, or paid for;
Whether it is normal access, or continuous, bulk, and high-frequency automated crawling;
Whether it bypasses captchas, access permissions, or other technical measures;
Whether it violates the platform's public and reasonable access rules;
Whether it increases the platform's server burden or affects normal operations;
Whether the acquired content is scattered information, or a data set with unique commercial value;
Whether the purpose after acquisition is to build indexes and assist retrieval, or to directly provide alternative services to the public.
Crawling public web pages to provide users with links to the original content, and extracting platform content in bulk to deliver it fully in another product, may receive completely different legal evaluations.
Therefore, crawling itself is just a technical action.
What truly determines the nature of the behavior are three questions: in what way the crawling is done, what is crawled, and what is done with the crawled content afterwards.
4. Not copying the original text word for word does not mean there is no problem
The unique feature of LLMs is that they do not necessarily present the crawled content to users exactly as it was. They can summarize, refine, reorganize, and rewrite it. This makes the dispute even more complex.
Copyright law protects original expressions, not facts, knowledge, ideas, and opinions themselves. If an LLM only extracts facts or opinions and then expresses them in new language, it may not constitute copyright infringement of a specific work.
However, if the model can stably reproduce important expressions from the original text, or if users only need to change a few prompts to obtain content that originally required visiting the platform or even paying to read, the situation will be different.
More importantly, copyright is not the only standard for judging such behaviors.
Even if the model's output is not substantially similar to the original text, another problem may still exist:
Does the LLM use the content resources formed by the platform's long-term investment through improper means, and create a substantial substitute for the original platform's services?
This is where anti-unfair competition law may come into play.
Copyright law mainly answers whether protected expressions have been used.
Anti-unfair competition law may further ask:
In what way was the content obtained;
Whether the platform has invested a lot of costs in building the content collection;
Whether the LLM company has saved the operating costs that it should have otherwise invested;
Whether there is an actual competitive relationship between the two parties;
Whether the LLM service reduces the necessity for users to visit the original platform;
Whether this kind of utilization goes beyond the normal and reasonable market boundaries.
Therefore, "the content has been rewritten" does not inherently mean it is legal, and "no verbatim copying" does not mean the dispute is over.
5. Not diverting traffic is not an illegal act in itself
Content platforms cannot require LLMs to take responsibility just because users choose to use LLMs instead of visiting their platforms.
Market competition inherently means that new products may replace old ones. Cars replacing horse-drawn carriages, instant messaging replacing SMS, and short videos changing text-based reading habits — none of these can be deemed illegal for new entrants simply because of "user loss".
LLMs do not have an inherent legal obligation to send all users back to content platforms.
Therefore, what the case really needs to examine is not "whether the platform has lost users", but whether the way the LLM gains its competitive advantage is legitimate.
This article is from the WeChat official account "Zhichanli" (ID: zhichanli), authored by Shawn/MCP, and published by 36Kr with authorization.