The intern who was fired by ByteDance raised 200 million RMB in financing, challenging Li Feifei.
The company does not even have a name yet, no product prototype is in sight, and the team only has 10 members, but it has already secured nearly 30 million US dollars (about 200 million RMB) in financing, with a post-money valuation of 200 million US dollars.
According to the latest reports, this sum of capital is invested in the world model company founded by Tian Keyu, with participation from 5Y Capital, IDG Capital and other institutions.
The most well-known label of Tian Keyu is "the intern who was dismissed and claimed for compensation by ByteDance".
How did an intern blacklisted by a large tech company make a comeback with a 200 million US dollar valuation two years later? And why is the "world model" he is working on so attractive that capital rushes to invest without hesitation?
One single paper drastically reduces costs for the whole industry
Tian Keyu graduated from the School of Software of Beihang University with a bachelor's degree, then entered Peking University for postgraduate study, supervised by Professor Wang Liwei, with research direction in deep learning optimization and generative models. Since 2021, he has interned in the commercial technology department of ByteDance, working on hyperparameter optimization, reinforcement learning and visual generation.
From June to July 2024, due to dissatisfaction with the internal computing power resource allocation of the team, he shut down the program of another intern to free up the graphics card for the research he recognized. ByteDance stated that he wrote and tampered with codes to maliciously attack the model training tasks of the team's research projects, resulting in resource loss.
After the incident, Tian Keyu once denied his behavior, claiming that the operation was done by another intern, and even called the police saying he was being slandered. ByteDance believed that he showed no intention of admitting his mistake, and officially sued him in November 2024, claiming 8 million RMB in compensation.
Later, he admitted his improper behavior. The court finally ruled that he violated the internship contract, ordered him to pay 500,000 RMB in compensation, and confiscated all his internship salaries.
A dramatic turning point occurred on the 10th day after the lawsuit was filed. In December 2024, NeurIPS awarded its annual Best Paper Award to the achievement with Tian Keyu as the first author — "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction" (VAR), which ranked 6th among all papers of that session. This is an extremely rare case for a scholar with Chinese mainland background as the first author to win this award.
The weight of this paper directly determines what he is doing today.
Before VAR, AI image generation mainly relied on diffusion models: first outline a blurry global picture, then refine it layer by layer, which produces good images but is slow, expensive, consumes huge computing power, and its architecture is not compatible with that of large language models. Tian Keyu and his team adopted a different approach: instead of flattening the image into a one-dimensional sequence for pixel-by-pixel prediction, they retained the two-dimensional structure, starting from a 1×1 blurry outline, then 2×2, 4×4, 8×8... predicting from coarse to fine level by level, just like how humans draw.
The result is that this GPT-style autoregressive model surpassed diffusion models in image generation for the first time: on the ImageNet benchmark, the FID (which measures distortion) dropped from 18.65 to 1.73, the inference speed increased by 20 times, and for the first time, visual generation realized the scaling law that "the larger the model, the better the effect". The open-source project has received more than 4400 stars on GitHub.
What Tian Keyu is best at is using a more cost-effective architecture to cut the cost of visual generation. Two years later, when he started his own business, he is still doing the same thing, except that the generation target is changed from images to videos, and from image generation to world models.
World model is at the peak of the trend
In simple terms, a world model enables AI to not only understand a single frame of image, but also perceive the space, motion and causality of the real world, and predict "what will happen next". It is regarded as the common underlying capability for robotics, autonomous driving and interactive videos, and it is also the direction that people like Li Fei-Fei and Yann LeCun have been pushing forward with all their efforts in recent years.
Tian Keyu believes that human language is fundamentally not suitable for describing videos. For a few minutes of footage showing a cat jumping off a table, most of the information including the position of objects, motion trajectory, light and shadow, and collision will be lost if it is written down in text, which also wastes computing power for no reason.
His solution is to create another set of language for AI. The team has already developed a "visual dictionary" containing about 200,000 symbols. These symbols are unrecognizable to humans, but AI can use them to represent, compress and predict videos. "We are first creating a language for AI, and then inputting 100 million hours of videos into it," Tian Keyu said.
With this method, the cost of generating 1 second of video is reduced to at least one tenth of the original level, and the full model is scheduled to be released in 2027.
The timing he chose is exactly the most bustling and subtle moment of this track.
On September 1, Li Fei-Fei's World Labs released Atlas, which can reconstruct 3D scenes from images, specify camera movement paths, and generate videos of up to 1 minute in length with 1440p resolution. On September 22, domestic company PixVerse released R2, which focuses on long-term memory, allowing users to change the world being generated in real time through text, voice and actions, with the plot generated on the spot. Overseas, Google DeepMind's Genie 3 already allows people to explore virtual 3D environments in real time.
On October 7, Google and Unity jointly announced Playground and the upcoming Unity Spark, enabling users to create and share small games using natural language.
Of course, the most shocking news for the industry is that on September 28, AMD suddenly announced that it would acquire World Labs for about 8.2 billion US dollars in an all-stock transaction.
Why did a company with sufficient capital, strong team and newly released models choose to be sold so quickly? Li Fei-Fei explained that the next generation of AI requires close collaboration of model research, systems and chips, and joining AMD can bring strong engineering and hardware capabilities.
This incident sends two signals to Tian Keyu. On the one hand, it points out a clear exit path for investors in the world model track, and re-ignites the valuation of this sector. On the other hand, even Li Fei-Fei has admitted that relying solely on the model team is no longer enough, and the key to success is shifting to computing power, cost and engineering systems — which is exactly why Tian Keyu bet on "cost reduction by an order of magnitude". In other words, he uses a 10-person team to pursue the same answer that AMD spent 8.2 billion dollars to acquire.
Why are 5Y Capital and IDG Capital willing to invest in a 10-person company with no products? One possible answer is that Tian Keyu is trying to solve the most expensive problem of world models: the computing power cost.
World models start to generate revenue
At present, the scenario closest to end customers is 3D content production.
World Labs' Marble has launched a monthly charging system: the standard version is 20 US dollars per month, the professional version is 35 US dollars per month, and the highest tier is 95 US dollars per month, with API and enterprise cooperation access provided. Users can input text or photos to generate explorable 3D spaces, export scene files, and then import them into design or game tools for further processing. The new model Atlas released in September further demonstrates the capability to generate 1440p videos of up to one minute in length according to specified camera trajectories, though the full model is still in the early access stage.
World models solve the most annoying rework problem for advertising agencies, film and television studios and game developers. The director thinks the camera position is too high and wants to lower it by 20 cm; the art staff wants to move the sofa to another position in the room. If each adjustment requires re-generation and re-layout, the efficiency brought by AI will be offset. Only when the same scene can be modified, exported and imported into the original production workflow repeatedly, can the related functions become part of the software budget.
In August this year, World Labs disclosed that production company Promise reconstructed ten historical crime scenes for the program "Harlan Coben’s Final Twist" on Paramount+ and CBS. Some of the locations no longer exist, and the team only had old photos or historical videos in hand.
Promise first restored the images, then used Marble to generate 3D environments, handed them over to art staff to fill in the missing parts and correct details, and finally placed the scenes in the LED studio for shooting.
Creative platform Magnific has also integrated Marble into its 3D Scenes tool. Marketers can turn reference images into 3D environments, place products in them, adjust camera positions and capture footage. When an advertisement requires multiple shooting angles, the team can continue shooting around the same scene without re-describing and re-generating the scene every time.
For investors, the larger market lies in factories, logistics and autonomous driving. If robots can complete a large number of training and tests in the digital world, the cost of trial and error on expensive real equipment can be greatly reduced. This demand is not imagined out of thin air.
In March this year, ABB Robotics announced that it would integrate NVIDIA Omniverse capabilities into its RobotStudio industrial software, and Foxconn is piloting related technologies for precision assembly of consumer electronics. ABB's stated goal is to reduce deployment costs by up to 40% and shorten product time-to-market by 50%.
Autonomous vehicles need to test scenarios such as rain and snow weather, road construction and sudden obstacles. It is slow to encounter these scenarios by driving real vehicles on roads, and it is also difficult to reproduce the same situation repeatedly. For robots to enter different families, warehouses and shops, they also need a large number of environment verification capabilities.
The opportunity brought by world models is to generate these environments in batches, so that enterprises can move more tests into the digital world.
On February 6, Waymo released a world model based on Google DeepMind Genie 3. The company disclosed at that time that its vehicles had completed nearly 200 million miles of fully autonomous driving on real roads, while the driving mileage in virtual environments reached billions of miles.
The three tracks of content, robotics and autonomous driving all have huge market potential, but their risks are completely different. General-purpose models require continuous capital investment to expand the leading edge, application-oriented companies must build sales and delivery capabilities, and the acquisition outcome depends on the willingness of a small number of strategic buyers.
This article does not constitute any investment advice.
This article is from the WeChat official account "Pencil News", author: Pencil News, published with authorization from 36Kr.