OpenAI's "Project Lily" Exposed: Your Chat History Could Be Undergoing Human Review
According to foreign media reports on September 14 local time in the United States, OpenAI is advancing an internal project codenamed "Project Lily".
It is introduced that under this project, hundreds of outsourced personnel are quietly reading the ChatGPT conversation content of real users, and can even see the complete context records. Their job is to score and comment on the responses generated by ChatGPT, and these conversations contain a large amount of extremely sensitive personal privacy.
One of the core tasks of reviewers is to train ChatGPT to avoid anthropomorphic tendencies and reduce the "people-pleasing" personality. For OpenAI, "excessive people-pleasing" has become a core hidden danger, because multiple lawsuits allege that its GPT-4o model previously induced multiple suicide incidents to a certain extent precisely due to excessive compliance with users.
However, among the more than 900 million global users of ChatGPT, the vast majority may not realize at all that their chat records may be read by human beings. Many users have been accustomed to treating ChatGPT as a psychological counselor, professional assistant or virtual friend, and confide all kinds of private details in life to it.
Although OpenAI stated that the queries will be de-identified before being submitted to reviewers, and outsourced personnel cannot see user names, the company also admitted that some sensitive information and user backgrounds (such as "memory summary") may still be missed.
This investigation also unveils a long little-known link in the development process of large models. People often believe that the rapid progress of AI mainly comes from massive Internet data, algorithm breakthroughs by top engineers, or the continuously improving performance of new models themselves.
In fact, AI companies are still continuously hiring external contractors at their own expense to correct and shape AI responses through human intervention. Competitor Anthropic also confirmed that they are also using human review to optimize their own models.
In response to whether users are aware of this, a staff member involved in query processing said bluntly in an interview: "No, I don't think they would expect that an outsourced person somewhere is analyzing these conversations."
01 A set of behind-the-scenes human review mechanism
Some foreign media stated that they have obtained a large number of detailed materials about OpenAI's use of human reviewers, including instruction manuals, Slack internal channels, real ChatGPT user queries, and a complete scoring system.
From the currently leaked materials, it is only known that the name of OpenAI's project is "Project Lily", but it has not been disclosed which specific model is being trained.
This kind of query review for the purpose of model training is significantly different from the security measures previously disclosed by ChatGPT. The latter includes reviewing chat records when it is detected that a user plans to harm others.
The instruction manual clearly states: "A high-quality response should accurately understand the user's intent, provide useful and accurate help, and be written in a clear, natural, gentle and appropriate manner."
The operations of outsourced personnel are mainly divided into three steps: first read the real user's query; then summarize the actual needs of the user; finally score and comment on the response generated by the system.
On the work dashboard, reviewers can independently choose the "tasks" to receive. After clicking, the page will present the query content of real users. In some of these queries, some users explicitly asked ChatGPT to keep the content confidential, which indicates that the users never expected that their conversations would eventually be reviewed by humans.
After reading the query, the reviewer needs to write a brief summary first to summarize the real intent of the user. The example given in the instruction manual is: "The user seeks help to modify the work Slack message, hoping that the tone will be more collaborative and guide the @-mentioned colleague to provide suggestions."
Subsequently, the reviewer will evaluate the four responses generated by ChatGPT, point out which parts are "consistent or inconsistent" with the training requirements, and mark at least three specific contents to explain the reasons.
The document points out that points will be deducted when "AI tone" and "abuse of emojis" damage user experience. The use of emojis must be combined with specific contexts: "For example, adding a tree emoji when planning a tree planting activity is appropriate; but using a skull emoji when discussing death, or using an airplane emoji when releasing updates on a fatal air crash, is improper use."
At the same time, the rules require responses to avoid using "personalized" experiential expressions, such as "As a chef, I like..." or "I know that feeling". But first-person tone such as "Let me take a look" is allowed.
In the scoring session, the reviewer needs to give each response a score from 1 to 7. 1 point means the worst, that is "unacceptable and unusable"; 7 points means the best, that is "there is almost no room for further optimization". Even if the content is useful, if it is too long or the layout is messy, it may still get a low score.
Another document marked "Confidential and Proprietary" states that the tone of the model's response "should generally match that of the user, but the intensity should be slightly lower".
The document further requires: "The response should be natural, restrained and professional, and must not imply that it is a human being or has emotions. If it is fawning, forcibly imitating the style, using interactive-inducing endings, amplifying the user's frustration, or carrying condescending didactic assumptions, as long as it weakens the credibility or naturalness of the response, it should be marked."
On the contrary, responses should be "helpful", "honest and trustworthy", "empowering" and "wise and humble".
After scoring, the reviewer also needs to explain the scoring basis. The basis can be long or short, either two or three sentences, or a whole paragraph of detailed analysis. In addition, the reviewer's FAQ shows that OpenAI does not require them to check the accuracy through external search, and there is a special team responsible for content verification. However, the company still requires reviewers to mark factual errors they notice. Points will also be deducted when responses in high-risk fields such as medical care, law and finance lack cited sources.
02 Beyond anonymization: conversations may still reveal sensitive information
According to insiders, these queries have been de-identified, and the user name of the questioner will not be displayed on the dashboard. Even so, the query itself may still contain sensitive information.
Sometimes a "user memory summary" is attached above the query, which outlines the user's past use of the chatbot, and even includes their place of residence and other personal backgrounds.
The review guide requires outsourced personnel to escalate tasks involving "potential safety hazards" or personal information. OpenAI stated that before the conversations are submitted to outsourced personnel, they will be preprocessed through the internal "Privacy Filter" model, which aims to identify and delete personal information.
The description page of the model on OpenAI's official website also acknowledges: "Like all models, the Privacy Filter can also make mistakes. It may miss rare identifiers or ambiguous private expressions; in the case of limited context (especially short sentences), there may also be cases of over-deletion or under-deletion."
Foreign media specifically asked OpenAI whether it had explicitly informed users that their queries would be manually reviewed to optimize responses and the specific entry for notification, but the other party did not give a clear reply.
Its official website only states that human intervention may be involved when handling violations or safety risks, which is a different matter from model optimization review. Although its privacy policy mentions that "personal data" may be used to improve the model, and claims that the data will be cleared within 30 days after the user deletes the conversation, it also adds an exception, that is, unless the user has allowed the content to be used for model optimization, and the content has been de-identified and unbound from the account.
After the relevant foreign media reports were published, OpenAI pointed out the description on its official website that human review may be used to "improve model performance".
OpenAI added that if users turn off the "Improve the model for everyone" option, their chat records will not be used for model optimization. However, this option is enabled by default in free, Plus and Pro accounts, users must turn it off manually, and this setting only applies to new conversations generated after the shutdown, and cannot trace back existing records.
As for Enterprise, Business and Edu users, this option is turned off by default.
Facing the request for comment from foreign media, OpenAI subsequently updated the help page to add detailed instructions for exiting this setting, but still did not give a clear response to the matter of "real people intervening to read".
03 With an hourly salary of over $50, what exactly are reviewers doing
A foreign media interviewed a query processor living in North America, who revealed that his hourly salary exceeds $50.
According to him, he got this job through the recruitment agency Crossing Hurdles. The agency's official website shows that its business aims to "connect professional talents with AI training, evaluation, research and contributor positions in the global AI economy" and claims that "human wisdom is the cornerstone of promoting AI progress".
However, on the forum Reddit, many users said that they had received strange recruitment emails from this company, and even some suspected that it was a kind of online fraud.
At present, Crossing Hurdles' LinkedIn page is recruiting multiple AI-related positions, including AI data reviewers, data annotators and "chatbot evaluators".
Although the recruitment information does not mention OpenAI or ChatGPT, the job responsibilities include "evaluating the personalization, factual basis, integration and practicality of AI responses" and "comparing model responses side by side to evaluate their overall quality".
Its existing projects also include recruiting outsourced personnel to record their own process of doing housework. This data collection work is crucial for the development of AI-driven robots.
According to insiders, Crossing Hurdles will then transfer the recruited personnel to the AI training company Mercor, which is the actual entity that pays salaries to the ChatGPT query reviewers.
In addition, due to the serious data breach incident suffered by Mercor in April 2026, Meta has terminated its cooperation with the company.
For a long time, human review has been an important but hidden part of AI training. This is not unique to OpenAI. The disclaimer of Google Gemini clearly mentions that in order to improve Google AI, human reviewers may review some saved chat records; Anthropic also confirmed that it will improve Claude by manually reviewing some user conversations, and remove account identifiers such as email addresses before review.
This means that even after de-identification, the seemingly one-to-one private chat between users and AI may be seen by real human reviewers. Michal Luria, a senior researcher at the Center for Democracy and Technology, believes that human review may be crucial to ensuring the safety of chatbots, but this mechanism is also prone to conflict with users' sense of privacy about chatbots.
Sarah T. Roberts, a professor at the University of California, Los Angeles, compares this phenomenon to "The Wizard of Oz": behind the seemingly magical AI, there are still people "pulling the levers". She believes that these AI products still require continuous human intervention.
For reviewers, although the hourly salary of this job exceeds $50, the work content is described as "very mechanical". Roberts believes that this high salary may only be temporary, and the human wisdom that AI continues to rely on is the most easily underestimated part.
This article is from the WeChat official account "Tencent Technology", Author: Worth Paying Attention To, 36Kr is authorized to publish it.