HomeArticle

Analysis on Legal Issues of "Dual Permission" for Artificial Intelligence Agents Operating on Behalf of Clients

互联网法律评论2026-09-04 07:59
Legal Analysis on Whether Third-Party Permission is Required for Artificial Intelligence Agents to Operate on Behalf of Customers

With the development of artificial intelligence technology, it has become increasingly common for AI agents to perform interactive operations on third-party apps, web pages and other platforms on behalf of users. After obtaining explicit authorization from users, AI agents performing online interactive operations such as account login, product browsing, form filling, and order query on third-party apps or web pages on behalf of users obviously helps improve the quality and efficiency of users' work and life. At the same time, some argue that the behavior of AI agents operating on behalf of users may also bring certain risks in terms of data security and privacy protection to users or third-party apps or web pages (hereinafter referred to as "third-party online platforms"). Therefore, whether the behavior of AI agents operating on behalf of users needs to obtain not only the explicit permission of the user, but also the "permission" of the third-party owner of the operated app or web page, the so-called "dual permission", is a key issue in current practice. Scientific research and analysis on this issue is of great significance for the healthy development of AI agent operation technologies and services on behalf of users.

I. Technical Paths of AI Agents Operating on Behalf of Users

At present, there are two main technical paths for AI agents to operate on behalf of users: one is the API (Application Programming Interface) invocation mode (hereinafter referred to as "API mode"); the other is the GUI (Graphical User Interface) simulated operation mode (hereinafter referred to as "GUI mode").

The API mode refers to that the operator of a third-party app or web platform actively opens standardized data interaction interfaces (APIs) and formulates invocation permissions, data scope, invocation frequency, and identity verification rules. The AI agent service provider obtains the API invocation key of the third-party platform through developer qualification certification, signing cooperation agreements and other methods. According to the user's natural language instructions, the AI agent service provider directly sends API interface requests to the third-party platform server to retrieve user data, issue operation instructions and complete operations on behalf of users. The API mode does not need to rely on the user's local browser to perform GUI simulated operations.

The API operation process mainly includes the following steps:

(1) The operator of the third-party app or web platform opens official API interfaces such as e-commerce order placement, order query, and product search to the public, and releases developer access rules and data usage restrictions;

(2) The AI agent service provider submits enterprise qualifications and business scenario descriptions, signs an API interface cooperation agreement with the third-party platform, and obtains exclusive API keys, OAuth and other authorization channels;

(3) The user grants authorization to the AI agent, jumps to the platform's official OAuth authorization page, and independently checks the permission scope allowed for the AI agent to retrieve (such as only viewing orders / allowing order placement / reading delivery addresses, etc.);

(4) The server of the AI agent service provider carries the key and user authorization credentials, directly establishes two-way data communication with the third-party platform server, and completes operations such as query, order placement, and payment on behalf of users through API interface instructions.

The API mode has the following characteristics:

First, the cloud server of the AI agent can establish a long-term two-way direct connection link with the third-party platform server, which can be completely separated from the user's local device;

Second, when the AI agent provides operation services on behalf of users to users, it can directly retrieve the underlying structured database data of the third-party platform, and obtain user privacy and transaction data in batches;

Third, the AI agent can directly interact with the third-party platform through exclusive developer keys, OAuth identity authentication and other methods. The third-party platform can accurately identify the identity of the AI agent service provider, and directly control and interfere with the agent operation services provided by the AI agent service provider.

Under the API mode, the third-party app or web platform must provide technical cooperation or collaboration for the AI agent, so that the AI agent can actually perform operations on behalf of users. Therefore, even if there are no restrictions in laws or contracts, the operation of AI agents on behalf of users through the API mode not only needs to obtain the permission of the user, but also needs to obtain the API authorization permission of the third-party app or web platform, which objectively obtains "dual permission".

If a large internet company group that owns influential internet platforms or apps in the market also intends to develop the AI agent operation service business on behalf of users, it will be more willing to adopt the API mode. These internet giants can quickly form an ecological closed loop of artificial intelligence and internet platform services. Therefore, the API mode is one of the solutions for some large internet companies to lay out the artificial intelligence industry at present.

The GUI mode refers to that the AI agent is deployed and runs on the user's local terminal device (such as a mobile phone, computer and other information processing systems). Through technologies such as screenshot, image recognition, OCR text analysis, and interface element positioning, the AI agent reads the information of the third-party app or web page displayed on the user's device screen, simulates the operations of human fingers, mouse, etc., to complete interactive behaviors such as clicking, inputting, sliding, and order placement; but the AI agent and its server do not establish any direct communication link with the third-party app or website server.

Take the case of Amazon.com Services, LLC v. Perplexity AI, Inc. as an example, the technical route adopted by the defendant Perplexity AI, Inc. for the intelligent agent operation service on behalf of users provided to users is a typical GUI mode. The complete operation process of the defendant providing operation services on behalf of users in this case is mainly divided into five steps:

(1) The user uses the Comet browser (provided by the defendant) on his mobile phone or computer terminal to log in to the Amazon account. The account password and session Cookie are all stored in the user's local device, and the Amazon server pushes page HTML, image and text data to the user terminal;

(2) The user issues natural language instructions to the AI agent (provided by the defendant) (such as "Help me select sports shoes within 300 US dollars and check out");

(3) The AI agent captures a screenshot of the current browser screen, and uploads the screenshot image data to the defendant's cloud server. The cloud server is responsible for analyzing page elements and generating next-step operation instructions, but the defendant's cloud server does not directly send any network requests to the Amazon server throughout the process;

(4) The defendant's cloud server generates an operation sequence and sends it back to the user's local browser, and the user's local browser program initiates standard user requests such as clicking, jumping, and order placement to the Amazon server;

(5) The Amazon server returns the operation result to the user terminal, the page is rendered and displayed in real time, and the intelligent agent takes screenshots again to execute the task in a loop until all the user's instructions are completed.

To sum up, the GUI mode does not need to negotiate interface cooperation with various apps or online platforms, and has the characteristics and advantages of low development threshold and full-platform adaptation. This mode has also become one of the technical paths for innovative AI service providers. Therefore, under the current exploration of intelligent agent technical solutions, a parallel scheme of the API technical model and the GUI technical mode has been formed.

II. Legal Nature of AI Agents Operating on Behalf of Users in GUI Mode

Although AI agents operating on behalf of users in GUI mode do not require technical cooperation or collaboration from third-party apps or web platforms, that is, only after obtaining the user's authorization, the AI agent can automatically and normally operate third-party apps or web pages on behalf of the user, third-party platforms, from the perspective of competitive interests, often claim that the AI agent still needs to obtain the permission of the third-party platform to perform operations on behalf of users. Whether the claim of the third-party platform is reasonable requires specific analysis of the legal nature of AI agents operating on behalf of users in GUI mode.

For third-party platforms, the operation behavior of AI agents in GUI mode on behalf of users belongs to the personal behavior of the user, not the behavior of the AI agent provider. Compared with typical civil entrustment, the operation behavior of AI agents in GUI mode on behalf of users has both similarities and differences. A typical civil entrustment is an act in which the trustee handles the entrusted affairs of the principal according to the entrustment of the principal. The similarity between a typical civil entrustment and the operation behavior of AI agents in GUI mode on behalf of users is that both are that the principal (user) uses the resources of a third party to handle civil affairs. The difference between the two is: a typical civil entrustment is that the principal uses a third-party civil subject to handle civil affairs; while the operation of AI agents in GUI mode on behalf of users is that the principal (user) uses a third-party tool to handle civil affairs.

Although some views or claims hold that the operation behavior of AI agents in GUI mode on behalf of users belongs to the behavior of the AI agent provider which is different from the user, such views or claims have been denied by judicial practice. In the case of Amazon v. Perplexity AI, Inc., the U.S. Court of Appeals for the Ninth Circuit explicitly denied the civil entrustment and agency claim of the intelligent agent. The plaintiff in this case claimed that the user's access to the plaintiff's web page, real-time interaction with the plaintiff's platform, and automatic shopping behavior through the AI agent provided by the defendant in GUI mode are behaviors attributable to the defendant itself, not the user's behavior. The U.S. Court of Appeals for the Ninth Circuit held that the user's behavior of accessing the plaintiff's platform web page through the AI agent provided by the defendant does not constitute the defendant's legal "access" to the plaintiff's platform system. In the legal sense or in the computer field, "access" refers to the act of entering the computer system itself or a specific part of the computer system (such as files, folders or databases). In this case, the defendant only provides the AI agent, which is a tool to help the user access the plaintiff's platform system, and it is the "user" who uses this tool to "access" the plaintiff's platform system. In this process, although the defendant's server will receive the screenshots forwarded by the AI agent or send instructions through the intelligent agent, these behaviors of the defendant cannot be regarded as the defendant's behavior of "accessing" the plaintiff's platform system.

It can be seen that the U.S. court regards the GUI mode AI agent operation service on behalf of users as a service tool, and the user's use of this service to perform network operations belongs to the user's own personal behavior, not the behavior of the AI agent provider. The above view of the U.S. court conforms to the practice of the GUI mode AI agent operation technology on behalf of users, which is worthy of reference for China. This is because the operation behavior of AI agents in GUI mode on behalf of users does have essential differences from the typical entrustment and agency behavior. In the entrustment relationship, the trustees are all subjects with independent judgment ability and civil liability ability.

In the operation of intelligent agents on behalf of users, although we can think that the intelligent agent has "independent judgment ability" to a certain extent, the intelligent agent itself obviously does not have civil liability ability. Although the intelligent agent provider can "execute and judge" the affairs entrusted to the intelligent agent by the user in the operation on behalf of users, this "independent execution and judgment" is initiated by the user's instructions and completed by the intelligent agent specifically, and the intelligent agent provider does not intervene in this instruction. Therefore, under the GUI mode, neither the intelligent agent itself nor the intelligent agent provider can be called the "trustee" of the user. In this case, it is most appropriate to treat the intelligent agent as the user's tool.

Since the intelligent agent is a tool for the user to perform automatic shopping and other operations on behalf of users, and the intelligent agent operates on behalf of the user rather than the intelligent agent provider, in general, the use permission authorization obtained by the user for using third-party apps or web pages can naturally be directly transferred to the intelligent agent service itself. That is, users do not need to obtain "additional permission" from third-party apps or web platforms when using the operation service on behalf of users in GUI mode. Of course, since the AI agent provider is not a party to the transaction of operation on behalf of users, the AI agent provider does not need to obtain the "permission" of the third-party app or web platform when the user performs the operation service on behalf of users in GUI mode.

III. Legal Validity of Contract Clauses of Third-Party Platforms Restricting Intelligent Agents Operating on Behalf of Users

When users use intelligent agents to automatically perform operations on behalf of users on third-party apps or web platforms, while facilitating users, it will also impact the core business models such as monetization of advertisements and traffic operation built by third-party platforms relying on manual interaction. Therefore, many large internet platform companies set up special clauses to restrict or prohibit users from using tools such as intelligent agents to perform operations on behalf of users through user service agreements, platform rules, privacy policies and other forms, via the path of contract law. The legal validity of such restrictive clauses is highly controversial in judicial practice.

Supporters believe that as an online service operator, the platform has the right to control its own service system, operation order and commercial resources, and regulating user behavior through contract clauses belongs to the legitimate category of business autonomy; opponents argue that such clauses are mostly standard clauses unilaterally drafted by the platform, which excessively restrict users' digital usage rights and independent choice rights, monopolize the platform traffic entrance in disguise, deprive users of the legitimate right to improve usage efficiency relying on new technologies, and belong to unreasonable restrictive agreements, which should be deemed invalid.

From the perspective of legal attributes, the contract clauses of third-party platforms restricting users from using intelligent agents to operate on behalf of users are typical standard restrictive clauses of online services, attached to the online service contract relationship between the platform and users, and are an important basis for the platform to regulate user behavior and delimit service boundaries. These clauses have the characteristics of unilateralism, standardization and universality. The platform pre-drafts the content of the clauses, and users can only choose to agree or reject the whole, with no room for negotiation and modification, which conforms to the legal definition of standard clauses in Article 496 of the Civil Code of the People's Republic of China. Therefore, the validity judgment needs to be based on the regulatory rules for standard clauses.

According to the validity rules of standard clauses in the Civil Code and the judgment logic of relevant judicial practices, the validity of the restrictive clauses of third-party platforms on intelligent agents operating on behalf of users is neither absolutely valid nor absolutely invalid, but needs to be comprehensively determined in combination with the content of the clauses, applicable scenarios, restriction degree and legitimacy basis.

The provisions of the Civil Code stipulate that the standard clauses shall take effect under two conditions: first, the provider of the standard clauses shall perform the obligation of prompting and explanation, the provider shall take reasonable measures to prompt the other party to pay attention to the clauses that have a significant interest relationship with the other party, such as clauses that exempt or mitigate its responsibilities, and explain the clauses to the other party as required; second, the content of the clauses does not have any legal invalid circumstances, and there shall be no circumstances that exclude the main rights of the other party, aggravate the responsibilities of the other party, or exempt the legal responsibilities of oneself.

In practice, some restrictive clauses provided by some platforms have procedural defects. Some platforms nest the restriction clauses on the use of intelligent agents in lengthy user service agreements, with no difference in font and typesetting from ordinary clauses, and do not perform the prompting obligation through prominent ways such as bolding, pop-up prompts, and separate confirmation, which makes it difficult for ordinary users to notice the restrictive agreement. In the era of artificial intelligence, prohibiting or restricting users from using artificial intelligence services has a significant interest relationship with users. Therefore, if the third-party platform fails to perform the statutory obligation of prompting and explanation, such clauses shall not have legal binding force on users.

At the same time, the rationality of the content of such restrictive clauses is the core key to the judgment of legal validity.

If such clauses only reasonably restrict the operation of malicious and non-compliant intelligent agents and prohibit behaviors that damage the platform order such as malicious order brushing and batch intrusion, they belong to the reasonable governance category of the platform. Such clauses do not exclude the main rights of users, and can be deemed legal and valid.

However, if such clauses uniformly and completely prohibit all behaviors of intelligent agents operating on behalf of users, such as users' independent authorization, small-scale, non-profit, non-malicious daily convenient operations, or users' reasonable operation behaviors on behalf of users through the GUI mode, they are excessive restrictions on users' rights and exclude users' main usage rights, which conform to the invalid circumstances of standard clauses stipulated in Article 497 of the Civil Code, and shall be deemed as invalid clauses.

This article is from the WeChat official account "Internet Law Review", the author is Yin Fenglin, and 36Kr is authorized to republish it.