首页文章详情

For the sake of your data security, do not use large language models as cloud storage.

三易生活2026-08-10 08:26
Large language models are not cloud drives, and you must take full responsibility for your own data security.

In recent days, some long-time ChatGPT users may feel a bit annoyed. Because OpenAI released an announcement at the end of July, stating that it will shut down the DALL·E image generation model entry in ChatGPT on August 30, 2026, and replace it with the new ChatGPT Images.

The most critical issue is that with the shutdown of DALL·E, all images generated with it in the past will be deleted from the servers. Therefore, OpenAI announced that it has reserved a one-month grace period. If users do not download their previously created images to local storage within this period, they will lose these data forever.

Why the deletion? Because the old and new models are not interoperable

First of all, it should be noted that OpenAI actually stopped using the DALL·E model as early as early May, and only retained its data access entry in ChatGPT. Although users can no longer use DALL·E to generate images, they could still view the previously generated images before.

With OpenAI's announcement to shut down the entry and delete data, it has actually given users more than three months of preparation time, which is by no means a "hasty decision".

The common "AI painting" style of DALL·E

But some people may wonder, when OpenAI takes offline the old model and launches the new one, can't it directly migrate users' past operation records and image data to the new system?

ChatGPT has actually reminded users very early that DALL·E is already "outdated"

Actually it can't. This is mainly because DALL·E is an "external plug-in" model. In the past, when users called DALL·E in ChatGPT to generate images, ChatGPT itself as the main model would first understand the user's intention, and then send the prompts of "what to draw and how to draw" to DALL·E.

In other words, DALL·E never knows why the user wants to draw a certain image from beginning to end. It only passively accepts the instruction of "what to draw" from ChatGPT. The data stored in DALL·E does not include the full context of the user's creative needs at that time, but only the prompts refined by the main ChatGPT model and the final images generated based on these prompts.

However, the new ChatGPT Images mode is not an external plug-in model, but a native multimodal model. That means it can understand users' conversations on its own, analyze "what to draw", and then create accordingly. This leads to the data format and context logic of ChatGPT Images being completely different from those of DALL·E. Even if the data of the old model is forcibly migrated to the servers of the new model, the new model cannot understand what these data without contextual conversations and only with prompts are for, and cannot achieve the purpose of tracing users' past conversations and creation history.

Therefore, directly deleting the historical data of the old model has become the only choice for OpenAI.

Cloud storage is not 100% safe, why not save data to local devices by default?

If you have used various AI applications before, you will find that the data such as conversations, images and audio in these applications will not be automatically saved to the local device by default after generation. Therefore, data loss caused by model updates or old model decommissioning like what OpenAI is doing is actually not uncommon.

Then why do all manufacturers not save data to users' local devices by default? In fact, there are several reasons for this.

First, it is not conducive to cross-device services. As we all know, today's large models can basically provide continuous cross-terminal services. For example, a conversation started on a mobile phone can be continuously viewed and "followed up" on a computer or tablet. Behind this cross-device capability, the biggest reason is that the large model will save the context and all kinds of generated data in the cloud by default. If the large model only saves data to the local device every time it runs, users will have to upload all the previous data first every time they "follow up a question" or use the service across devices, which will significantly increase the waiting time.

Second, the cloud data storage mode also makes it possible for large models to be deployed on devices with different performance levels. For example, some feature phones and smart wearable devices with low computing power can now access large model functions. But if all data is automatically saved locally, these devices with extremely limited built-in storage space will soon be fully occupied, which will cause inconvenience in use.

In addition, the default cloud storage mode also lays the foundation for compliance management of large model conversations. After all, if the data is only saved on the user's device, it will be difficult for large model service providers to find non-compliant content in the first time. When an accident occurs, since users may modify or delete local files, the original conversation records saved in the cloud will also become the basis for distinguishing rights and responsibilities.

Finally, even for "data security" itself, cloud storage may help service providers avoid liabilities more easily. After all, if data is saved locally by default, once a user accidentally deletes the data and refuses to admit it, it will be very difficult to clarify the situation. On the contrary, the cloud data storage mode widely adopted by AI service providers now has clear terms: the platform only provides "online viewing" services and does not undertake the obligation of permanent storage. If users really care about their historical data, they should back it up by themselves.

Large models are not cloud disks, you should take responsibility for your own data security

In fact, from OpenAI's perspective, they have reserved more than three months for DALL·E users to back up and export their data. Considering that the storage capacity of relevant infrastructure during this period is occupied by old data that can no longer generate profits, it can even be said that OpenAI has done its utmost in this matter and there is nothing to criticize.

But from the users' perspective, the loss of conversation records, creation history, and even some of their "works" is a foregone conclusion. They certainly have reason to feel that they have been treated as victims of technological progress and business iteration.

So this incident tells us that although many people have regarded AI as a partner in work and study, and even pour out their troubles in daily life to it, AI technology is constantly evolving, and the existing models will eventually be iterated one day. Maybe at that time, the historical conversation data will disappear, and the familiar conversation style will also "become different". Therefore, if you really care about your own data, irregular backup will be very necessary.

After all, large models cannot be used as cloud disks, and AI manufacturers have never promised to accompany you forever.

This article is from the WeChat Official Account "Three Easy Life" (ID: IT-3eLife), author: San Yijun, published with authorization from 36Kr.