100,000 people are queuing up for access, with finished content generated in one click. The easier AI creation gets, why has it become more difficult than ever to create viral hits?
When AI can "create" from zero frames, content production has become an "excess" chain. Over the past year, multimodal models have pushed the technical threshold of content production to an unprecedented low. Writing scripts, creating storyboards, generating images, and even completing editing and dubbing can all be done by one person, on one computer, or even with a single sentence.
But a more critical question emerges: when "production" is no longer scarce, what is the new scarce capability?
Focusing on this question, we invited two frontline practitioners to share their answers in a roundtable dialogue during WAIC. In the conversation, Owen, Product Manager from OiiOii, and Bruce, Chief Operating Officer of SeaArt.ai, deconstructed the real variables behind AI creation from real user cases, product decisions and user ecosystem: how the threshold disappears and where opportunities re-emerge.
The following is the sorted transcript of the dialogue:
Where exactly has the AI creation threshold dropped to?
InfoQ: We can now see that AI has lowered the creation threshold. Where exactly does the reduction lie? How low is the threshold? The intuitive experience for most people may be generating multimodal content with a piece of prompt, registering as a member, and spending points. But in addition to light use, what other changes and future directions has it brought to the industry?
Bruce: I can share a real user case from our platform. When doing user research before, a user left a particularly deep impression on me. He is from Australia and works as a barber. He dreamed of becoming a painter when he was a child, but he put down his brush and picked up scissors for a living. During the exclusive interview, he showed us his account and said that the happiest moment of every day now is to turn on SeaArt.ai after getting off work and create around his original OC character. There are hundreds of videos, thousands of pictures generated around his favorite characters, as well as corresponding models and applications in his account. He created a character and then kept creating around this character, which is like presenting the perfect image in his mind through the act of creation.
This case particularly illustrates the essence of AI lowering the creation threshold: in the past, capability was the boundary, but now imagination is the boundary. What we need to do is to keep pushing the boundary outward, so that more people who originally had no chance to express themselves can present the perfect world in their minds. Our vision is to enable everyone to have the ability to express themselves. Whether they are professional creators or ordinary users, we can provide creation tools that are easy to get started with but difficult to master thoroughly.
Owen: From my perspective, the reduction of AI creation threshold is mainly reflected in two dimensions. The first is the disappearance of skill and equipment thresholds. In the past, if ordinary people wanted to shoot videos, they needed professional equipment, and they needed to know how to write scripts, find scenes, find actors, shoot, edit and dub. With large models and multimodal models, the language model can take over all the scripts, character settings, scene and picture prompts, and the multimodal model finally converts these texts into pictures. For ordinary people, from the idea in their mind to the final visible work, there is no need for practical operation in the physical world anymore, which is the equalization of the right to use cameras.
The second is the cliff-like drop of cost, including trial and error cost and time cost. In the past, a small team might spend a whole day making a 60-second video. Now with multimodal tools, one person can complete the whole process and produce a high-quality finished film in maybe one hour. Judging from actual user cases, in the past six months, more and more non-professional creators with zero foundation in our ecosystem have created hit works with hundreds of thousands or even millions of likes on Douyin through OiiOii, and these cases are no longer isolated, emerging every week. These creators may be students or office workers in foreign trade companies during the day, and turn into producers and directors at night, which intuitively shows how AI video tools and image tools change the creation link of hit content.
InfoQ: Around this time last year, I communicated with practitioners in the art industry. He said that using AI to create paintings still requires mastering professional terms such as scene scheduling, composition, and lighting to instruct AI. At that time, I felt that the AI creation threshold had not been lowered very much, because people who did not understand the terms could only describe a very subjective paragraph, and the result depended entirely on how AI understood it. Combined with the cases described by the two just now, does this threshold still exist now? Is it necessary to know professional terms to make good works?
Owen: The concept of "prompt wizard" may have gradually disappeared. First of all, multimodal models can already infer the general picture you want directly from very plain descriptions of pictures, shots, and light effects, and the threshold has dropped a lot. Secondly, with the improvement of the writing ability of language models, you only need to purely describe what happened to the characters, including time, place, characters and plot, and it can help you convert it into content that only professional screenwriters can write. What we do for tools is to find ways to fully unlock the black box capabilities of multimodal models, so that ordinary users can also use 100% of the capabilities. You only need to clarify the main line of the story and handle the logical relationship well, and those special effect descriptions, complex camera movements, and shot reverse shots can be filled by the model, and you only need to do some fine-tuning.
Bruce: Regarding the issue of whether the last mile of AIGC landing in creation has been solved, my view is that the capabilities of language models are already very strong now. The fundamental concept of our product is to let everyone get started easily and also let professional creators stay, that is, to achieve the state of easy to use but difficult to master. When SeaArt started as an image community in 2023, Stable Diffusion had strong capabilities but its interface was like an aircraft cockpit, and it took a long time just to memorize more than 70 parameters, which we also spent a lot of time learning at that time. The opportunity for SeaArt to be born is to solve this last mile problem. We have made a one-click image generation tool, so that users can produce works beyond expectations even without any prior AI background knowledge. This is also the reason why we can develop well: popularize this technology and solve the last mile problem.
From Zero to Hit: How Are Demands "Seen"?
InfoQ: Both of your products are very successful and have achieved considerable market results in the target market. What was the opportunity for your products to enter the market at the very beginning? What opportunities did you see at that time? What was the group of your first batch of core users?
Bruce: Before we built SeaArt, we had a game company called Xinghe with more than 1,000 employees in Chengdu, a large part of whom were art staff responsible for 3D material production and game art. Before the emergence of AI, producing game art materials was time-consuming and labor-intensive. After tools such as Stable Diffusion and Midjourney appeared on the market in 2023, we were the first batch of companies in the industry to apply these tools to the art material production process, and the effect was immediate: game materials that used to take a week to produce can be completed in two days. But we also encountered problems in use. Stable Diffusion had too high a threshold, and even our very senior art staff needed to spend about 10 days figuring out how to use it. Although Midjourney is easy to use, it is a closed-source model and cannot meet our highly personalized material production needs.
At that time, we thought that since there was no perfect solution on the market, we might as well make one by ourselves. After it was developed, the internal response was very good: everyone reported that it was easy to use, and the output results were much better than the previous effect that could only be used after trying many times. We realized that since there was such demand internally, external creators must also have this demand. For example, a large number of people who have imagination but do not know how to turn their imagination into reality need this tool. So we officially launched it to the public in April 2023. Up to now, it has covered more than 50 million users in 192 countries around the world, and has become a national-level application in Japan.
InfoQ: When you opened this tool to the public at that time, weren't you afraid of empowering competitors in the game industry?
Bruce: If we can empower our peers, it is also our good karma, which can drive the whole industry to progress together. Later, we found that this is not only an opportunity at the tool level, but also an opportunity for the transformation of creation methods. AI today is like the camera in the mobile Internet era, which is a paradigm-level change. This is the opportunity for us to enter this market at the very beginning, and the root of our good development in the future.
InfoQ: In the past, game companies employed a large number of original painters and material painters. Now they still need to recruit people to use AI and be responsible for the final review of manuscripts to decide which version to use. What is the talent profile of this position now? Do you value ideas and aesthetics more, or do you still prefer people who transferred from the original painter industry to do AI painting?
Bruce: The talent recruitment profile has indeed changed. Before 2023, when we recruited people, we valued their past work experience, whether they could use game material production tools, and whether they could produce the required 3D materials. But now capability is no longer the boundary, and imagination is the boundary. When we recruit game-related talents now, we first look at whether this person has spirituality, that is, whether this person is interesting, and whether he or she has points that can resonate and bring happiness to others. In the content production industry, this refers to the ability to define hit works. From our past experience, this ability is very difficult to cultivate, it can only be screened, and it depends a lot on talent. The content industry seems to have a low threshold, but the subsequent climbing curve is very steep. It is very difficult to climb up without talent, because if you can't arouse the resonance of users, it is difficult to improve the long-term retention.
InfoQ: Owen, the breakout moment of your product is more recent. Which year was it at the very beginning?
Owen: Our breakout wave was in 2025. After images comes video. The opportunity for us to enter the market is that we see the models have become really usable, and the produced content has consumption value. Before Sora came out, models such as Keling could already produce consumable content through multiple attempts, and many short drama companies were also trying to use AI to reduce costs and increase efficiency. Although the models have become usable, for ordinary users, there is still a big gap from not knowing how to use the model to producing the final finished film, and then to making high-quality films, which is the opportunity for tools. This is the situation for ordinary users. For B-end customers, they need to better control assets and materials, reduce the rate of invalid attempts, and improve content conversion rate. The agent built into the tool has its own industry know-how, which can help users write prompts, increase the output ratio from 1 qualified work out of 10 attempts to 1 out of 5 or even 1 out of 3, so the cost and time are naturally reduced.
In addition, many people worry that the continuous iteration of model capabilities will eat up the value of tools, but we judge that this will not happen in the video field. In the long run, the value of tools will always exist and is difficult to be internalized by the model, because from the idea in mind to the consumable finished film with hit potential, it requires close combination of creators and tools, conception and fine-tuning. Tools will have great value for a long time.
InfoQ: Owen, do you watch manhua dramas? From your personal point of view, what is the production level of manhua dramas now?
Owen: I do watch manhua dramas, and I used to watch more male-oriented ones. Recently, many manhua dramas have begun to top the ranking list. At first, live-action dramas were popular, then hyper-realistic human dramas became popular, and recently manhua dramas have risen again. With the maturity of tools and model capabilities, content consumption forms are also changing. The production quality or production difficulty of manhua dramas may be slightly lower than that of hyper-realistic human dramas. The cost of hyper-realistic human dramas is higher, because to make the characters look extremely like real people, the emotions need to be delicate enough, with micro-expressions and micro-actions, which requires more prompt fine-tuning and attempts. Manhua dramas have an inherent advantage: when the audience sees this style, they already know it is generated, so they will not be easily distracted, and can focus more on the plot and script itself.
Hit the User's Satisfaction Point: Functions Voted by Real Money
InfoQ: Next question. When was the first time that both of your products achieved large growth due to a certain function? What was the function? How did you discover it? Let's start with Owen. The scene of 100,000 people queuing up happened right in front of us.
Owen: The breakout point at the end of last year was the function that focused on seven agents to help you make animations. After entering a few sentences in the tool, multiple agents such as storyboard designer and screenwriter will come to work, giving users the feeling of directing the whole team to work, which brings great impact and satisfaction, and is a typical aha moment. From the perspective of animation production process, before the emergence of AI, the threshold for making formal animations was extremely high. You had to spend half a year learning tools such as PS, Blender, and Maya, but that was just stepping into the threshold. Then you had to learn to write scripts, design storyboards, create characters and scenes, design key frames, do compositing and dubbing. Each link had an extremely high learning threshold, which was difficult for ordinary people to reach. Now the model has flattened all these thresholds, you can direct a group of digital employees to work for you, and build your own animation team.
For C-end users, daily life stories that they are usually unwilling to shoot or find it costly to produce can be generated into animations through AI and posted on social media at a low threshold. For B-end users, although the function of splitting long scripts and one-click localization for overseas markets was not supported at that time, it took them a long time to manually split scripts and write prompts. With the help of agents, the work can be done very quickly, and only one person needs to do the final check and adjustment. So this function is of great value for both C-end and B-end users.
InfoQ: Bruce, what changes did your product make that led to the large-scale growth of SeaArt at the very beginning?
Bruce: Our first wave of breakout happened at the time of launch, SeaArt became popular as soon as it was launched in April 2023. At that time, there was a volunteer in Brazil, who was not invited by us, but a real organic user. After he used SeaArt and found it very easy to use, as a big V blogger, he posted the usage process on YouTube, and SeaArt became popular in Brazil in the first wave. The obvious growth in 2024 came from the creator incentive plan we launched at the beginning of the year. In short, creators publish their self-created models and applications on the SeaArt platform, and the platform pays cash rewards directly according to the number of runs of their works. The response of this plan was very obvious: the content supply on the platform increased instantly, more users came to use the models and applications published by others, and the flywheel of production, distribution and consumption started to rotate.
The reason why we launched this plan is that we reviewed the successful content platforms in the mobile Internet era, including Douyin, Bilibili and Xiaohongshu, and found that they all built a complete business closed loop, where creators can get paid for producing high-quality content on the platform. If we always let creators produce content for love, creators will leave and the community will die. If we can let everyone make money on the platform, the driving force will change from initial love to real returns, and the platform will develop from a community to a place that brings both fame and wealth for creators.
InfoQ: This involves a very key issue: payment. All communities and C-end products cannot finally avoid the question of how much users are willing to pay for subscription. In 2023 and 2024, the widely discussed problem of AI products was: we spent so much money training large models, how much are users willing to pay for using them? At that time, people might try it for fun, but they were not willing to spend a lot of money on subscription. But today, there are a large number of people who spend one or two hundred dollars a month to subscribe to a single agent tool, and tens of dollars to subscribe to the AIGC platform. The key is that the value of a certain product is perceived by everyone, and they are willing to pay for it. What functions of SeaArt make users willing to pay? How do you evaluate this?
Bruce: For our positioning, SeaArt.ai is not just a tool, but an AI interactive entertainment platform. Now AI does have a trend of being abused because the threshold is low, but users will not pay for more content, they pay for the relationship between themselves and the characters. Just like that Australian barber, he created original characters on SeaArt, and built emotional connection with these characters, he is paying for this emotional connection. The bathroom in our SeaArt office is pasted with a slogan "We need to provide cost-effective emotional value", users are essentially paying for emotional value. Our principle