Claude and GPT are venturing across their original domains into the video sector, are the competitive moats of players such as Seedance still rock-solid?
Over the past month, the most exciting trending discussion in the AI video sector has not revolved around professional video models such as Seedance, Minimax and Keling, but general large models that have "crossed boundaries" into this field.
Claude Opus 5.5 renders videos through code, spawning a large number of popular science animations, creative short films and music MVs in just a few days. Multiple hit works have emerged on platforms such as X, even sparking a brand new creative paradigm: Vibe Video.
GPT-6 Astra acts as an "AI director", completing storyboard planning and white mold pre-visualization. By deeply integrating with professional software including Blender, Unreal Engine and DaVinci Resolve, it undertakes tasks such as scene construction, camera pre-visualization and editing execution, gradually taking up the upstream segment of video generation.
The two follow two completely distinct paths: one is programmatic creativity based on code rendering, the other is film and television industrialization based on spatial planning. Neither of them directly competes in image quality, but cuts in from upstream entrances and segmented scenarios, quietly rewriting the capability boundary of AI video.
While general large models are crossing over from the peripheral and quietly rewriting the capability boundary of AI video, the layout of professional video models is also accelerating. Leading players including ByteDance, MiniMax and Kuaishou are collectively penetrating into the upstream of creation, gradually covering the complete production chain of script decomposition, storyboard scheduling and finished video editing. This deep rooting in the entire creative process endows them with stronger irreplaceability amid this round of cross-border impact.
01
Opus 5.5 sparks the popularity of Vibe Video
Does code directly "draw" videos?
On September 22, Anthropic officially released Claude Opus 5.5. The official focused the release on programming capabilities and Agent task execution, and video capability was not the core highlight of the product launch. But in subsequent developer tests, an unexpected viral breakthrough quietly emerged: although this model does not directly output pixel images, it can generate animation videos with extremely high completion by writing front-end code.
This is not a brand new idea. Previous pure code drawing tests such as "SVG draws a pelican riding a bicycle" have long been circulated in the community. But the code stability and task planning capability of Opus 5.5 have pushed this creation method from toy-level demonstration to full implementation. Relevant works first emerged intensively on the X platform, and then spread to domestic platforms such as Xiaohongshu and Douyin.
Some creators joked that in the week after the release of Opus 5.5, their timeline was almost filled with all kinds of AI code videos.
Amid the boom, the most representative benchmark case that achieved widespread cross-circle popularity came from X platform blogger @anabology. He only provided Opus 5.5 with a reference prompt, Midjourney access permission and a mood board, and the model independently started the whole production process, outputting a dystopian short film with cinematic texture 12 hours later. Up to now, the number of views of posts related to this work has exceeded 20 million.
Elon Musk showed high attention to this work, reposted and interacted with it for many times, and said bluntly "This time I truly feel the far-reaching impact of AGI". He also replied "Accurate" to a post claiming that "Opus 5.5 has reached 80%-90% of the level of artificial general intelligence". The public recognition from direct competitors further pushed the video capability of Opus 5.5 to a wider audience.
A large number of highly discussed cases have also emerged in scenarios such as popular science short films and product promotion videos. Blogger vittorio only gave one creation instruction, and Opus 5.5 produced a 2-minute and 16-second short film on the history of Western civilization, without calling any professional video generation model. The video gained more than 10 million views in two days after release.
Some creators relied on Opus 5.5 to write about 7400 lines of front-end code based on the Remotion framework, automatically completed the arrangement of full-frame motion effects and transition rhythms, matched with open source speech and soundtrack generated by Python, and produced a 3-minute video themed on the history of AI development in only one hour.
Inference startup company Deedy used Opus 5.5 to make product launch videos. The generation time for a single video was only about 1 minute, the API cost was only about 2 US dollars, and it finally gained 320,000 views. The Wasmer team directly input the text content of the product launch announcement into Opus 5.5, and the model generated a complete brand promotion video at one time, whose communication effect exceeded the team's expectation.
These cases have quickly made "Vibe Video" a new creative paradigm: without relying on professional video models, you can produce complete videos only by using large models to write code. The case collection site awesome-opus5-5-videos on GitHub has included more than 400 works, covering categories such as motion graphics, knowledge explanation, 3D scenes and small games. Each work is attached with the original prompts and code made public by creators, which is equivalent to directly opening source the mature creation methodology. Domestic creators have also entered the market one after another, using it to make product promotion videos, popular science animations and data visualization videos.
The video generation logic of Opus 5.5 is essentially different from traditional video generation models.
Traditional models generate complete frames at one time based on massive training materials, while Claude follows the pure code rendering path, building frames frame by frame like a professional motion designer — defining all visual elements and motion trajectories by writing front-end codes such as SVG, Canvas, Three.js and Remotion, then running them in a headless browser and recording frame by frame, finally encoding and synthesizing MP4 files via FFmpeg. During the process, it will also automatically check frame images and correct deviations.
In the whole process, the large model plays four roles at the same time: acting as a director to finalize the storyboard rhythm and visual style, acting as a producer to decompose tasks and schedule multiple segments of code for parallel generation, acting as an artist to draw frame elements frame by frame with code, and acting as an editor to complete audio-video synchronization and finished video output.
The biggest advantage of this path is extremely strong editability and extremely low iteration cost. If you are not satisfied with the result of traditional video models, you often can only modify the prompt and regenerate the whole video; for code-generated videos, you only need to modify a few lines of code of the corresponding shot, the parameters, motion and style can be precisely controlled, and the frame elements can also be disassembled in layers.
But the boundary of application scenarios of Opus 5.5 is also very clear. It is naturally good at programmatic and graphic content, such as text motion effects, data charts, popular science animations and abstract 3D scenes, but it is difficult for it to handle realistic real-person frames and complex natural physical effects, which are exactly the core home ground of traditional video models.
This also determines that the current penetration scope of Vibe Video is still concentrated in the tracks of lightweight commercial content, popular science and creative experiments, and has not yet entered the mainstream market of film-level realistic content.
02
GPT-6 Astra takes over the director's chair
Will video models be reduced to "rendering workers"?
If Opus 5.5 cuts into creative videos from the code dimension, GPT-6 Astra extends its reach to the more upstream planning link, reshaping the production logic of AI videos. It also does not directly generate the final frames, but plays the role of "director + producer": first complete creative planning, spatial design and camera scheduling, and then hand it over to professional video models for rendering.
GPT-6 Astra was officially released on September 3, positioned by the official as a new generation of general agent model, whose core capabilities cover fields such as computer operation, programming and scientific research. Among them, its operation capability for 3D creation software such as Blender and Unreal Engine has become the key fulcrum for it to cut into video creation.
A test publicly released by Higgsfield AI is quite representative. The task is to make a one-shot animation that continuously traverses multiple spaces. Astra does not generate images directly, but first builds a white mold scene with simple geometries in Blender, plans the character's movement path and the whole camera movement route, and after confirming that the spatial logic and camera movement are correct, it hands it over to Seedance 2.5 for rendering into a finished video.
In the final finished video, the camera penetrates from one space to another, and the character position and scene relationship are always consistent, which alleviates the long-standing problems of spatial disorder and camera penetration in AI videos to a large extent.
The full-process creation capability of Astra is also quite remarkable. Overseas creator Nate Herk did an experiment: he only gave an open instruction to let Astra make a YouTube video about its own release. Astra independently completed data research and script writing, called tools to generate cloned voiceover, drove AI digital humans to appear on camera, completed material editing and music sound effect matching, and finally delivered a finished video that can be released directly.
The access of professional editing software further amplifies this capability. On September 8, DaVinci Resolve 21.1 launched a native MCP server, supporting AI assistants to directly operate the software through protocols. A 20-year experienced editor handed over 75 minutes of original materials to Astra, which directly operated the editing software and output a 22-minute rough cut version containing 96 editing clips, subtitles, color grading and volume calibration, without any manual intervention in the whole process.
Domestic creators quickly explored localized workflows. Some people handed over a complete Xianxia script to Astra to decompose storyboards and design shots, and then connected to Seedance 2.5 to generate a 90-second finished video, whose array design and shot switching completion exceeded expectations; other creators summed up a relatively mature product promotion video workflow: first GPT-6 does white mold pre-visualization to verify the feasibility of the shots, then export the pre-visualization video to the video model for rendering, and finally synthesize it with editing tools. The whole process takes about two hours, and the cost is only equivalent to 150 RMB.
What supports these performances is the world model and 3D spatial understanding capability of GPT-6 Astra. It can understand the topological relationship of multiple connected spaces, the internal logic of character movement and the motion path of the camera, rather than just generating flat images one by one.
The greater value is that creativity has thus become a modifiable and reusable asset: the spatial layout, character scheduling and camera routes can all be retained, so you don't have to start from scratch when replacing characters, scenes, art styles or even stories.
At this point, general large models have come up with two differentiated paths, forming an interleaved competition pattern with professional video models. The code faction represented by Opus 5.5 is good at programmatic creativity and motion effect popular science, with low cost and fast iteration, suitable for lightweight commercial content.
The planning faction represented by GPT-6 Astra is good at complex camera scheduling and spatial narrative. Paired with video models, it can produce film-level content, which is closer to the professional production process. The two are not in a substitution relationship, and can fully cooperate: GPT-6 is responsible for storyboard planning, and the video model is responsible for rendering the realistic part, forming a complementary workflow.
03
Video models collectively move upstream
Full-link layout reinforces the moat
The two cross-border paths are quietly reconstructing the profit distribution of the AI video industry.
On the one hand, commercial and creative content such as product promotion videos, popular science and data visualization are beginning to be diverted by Claude's low-cost code path. On the other hand, GPT is beginning to master the division of labor in video generation, reducing video models such as Seedance to downstream rendering processes, which may lower their value in the industrial chain in the long run.
While general large models are disrupting the track from the periphery, professional video model manufacturers such as Seedance, Keling and MiniMax are also reinforcing their moats at their own pace. All manufacturers share highly consistent directions: extend to the upstream creation link, deeply embed the complete workflow of creators, and replace the single-point material generation capability with end-to-end production capability.
On September 20, Jianying launched the one-stop creation workspace "Jianying Hub" and the intelligent creation Agent "Jianying Assistant" at the "AI New Creation Conference". Combined with the previously launched Skylark Agent and