HomeArticle

Many people are still unaware of the real megatrend of the next wave of AI.

美股投资网2026-09-09 11:03
It is very important!

I have been testing every major AI update in the past two years as soon as it is released. From Agent systems and multimodal capabilities to wave after wave of model iterations — to be honest, most of these updates no longer surprise me. Their performance is indeed improving, but they have all stayed on the track of "making generated content more realistic" without any fundamental shift in direction. It was not until this month, when I ran through Astra, Atlas and Cosmos in succession, that I stopped and truly took notice for the first time.

This is not because the content it generates looks more polished, but because for the first time I feel that AI is crossing another threshold — shifting from "generating content" to "understanding the world and getting ready to take real action". I will put forward this conclusion first: the next round of real value transfer will not lie in the competition for model parameter counts, but in the fact that the three-dimensional world is being rewritten into an interface that AI can read, simulate and operate. If you only understand this trend at the level of "there will be more demand for 3D content", you will miss the entire replacement of the underlying logic. This statement may sound overly assertive, but after testing these three technologies in full, I cannot convince myself that this is just a coincidence.

1/ Astra: The Cost Structure Collapses First

I asked GPT-6 Astra to generate a residential interior scene. In the past, no matter how good-looking the output of this kind was, it was essentially useless once imported into modeling software — all structural information was lost, and the output could not be used for subsequent work.

This time it is different. It directly outputs assets with complete structures: they can be further edited in Blender, and imported into Unreal Engine 5 for real-time roaming immediately.

Anyone who has worked with 3D workflows knows how expensive this process is: modeling, texturing, lighting, rendering, post-processing, every step requires manual operation in professional software, with extremely low fault tolerance rate and extremely high time cost.

Now? Natural language is the entry point.

Images and videos have already staged this exact story: once the cost of tools is driven down, demand will not be suppressed, but released. The gaming, e-commerce, architecture, advertising and industrial design industries are all queuing up for this. This is not a hypothetical deduction, but something that is happening right now.

2/ Atlas: From "Creating Objects" to "Understanding Space"

If Astra made me notice the change, then Atlas, released by World Labs in September, is what made me truly realize how far this round of development can go.

Think about this difference: when generating a cup, you only care if its appearance is correct; but to understand a kitchen, the model needs to know where the cup is placed, how high the countertop is, which direction the hand should extend from most naturally, and what will happen next in the space.

The former is creating objects, while the latter is understanding space. What separates them is not clarity, but an entire dimension. There is one sentence from World Labs that I have read over and over: "3D is becoming the universal interface for space."

Text used to be the interface between humans and machines. If 3D becomes the medium for humans and AI to jointly generate, edit and simulate space, then 3D will no longer be a single design file, but a new computing carrier. This judgment may still sound far-fetched, but from my actual testing experience, it is becoming increasingly concrete.

3/ Cosmos: From "Perceiving the World" to "Taking Practical Action"

When it comes to NVIDIA's Cosmos, I am basically sure that my previous feeling is not overinterpretation.

Current visual models can already accurately answer "what is in the picture". But the real difficulty for robots and autonomous driving has never been recognition, but judging "what will happen next".

What Cosmos does is to allow the model to first build an internal deduction system for world states — simulate the possible outcomes of each action, and then select the optimal path. It is equivalent to installing a runnable internal simulator for AI.

Without it, machines can only perceive the current moment; with it, machines gain the ability to "deduce the whole process before taking action". I believe the significance of this step to Physical AI is no less than the breakthrough in reasoning capabilities of large language models back then.

4/ The Real Game-Changer: Synthetic Data

Data from the real world is extremely expensive.

Language models can crawl massive amounts of text from web pages, but robots cannot learn how to grasp a cup from web content; autonomous driving cannot possibly cover all extreme scenarios through real road tests. And the most critical training data — extreme weather, unexpected accidents, rare road conditions — is precisely the hardest to obtain.

But once 3D reconstruction and world models are fully mature, the logic will change: first build a digital reference scene with limited real data, then systematically modify the weather, position, object attributes and lighting in the virtual environment — to generate massive training samples from a small set of seed data.

Therefore, analysis from the US Stock Investment Network believes that the competition dimension of Physical AI is shifting:

In the past, the competition focused on who owned more real data; in the future, the competition will center on the product of real data quality × spatial reconstruction capability × simulation accuracy × synthetic data amplification efficiency.

Beneficial U.S. Listed Companies

I divided them into four tiers according to "the directness of benefit": the higher the tier, the more certain the benefit, and the lower the tier, the greater the elasticity and the more concentrated the exposure: Tier 1 · Infrastructure: NVDA

No matter which robot company or autonomous driving company wins in the end, they all need computing power, simulation, synthetic data and training platforms. NVIDIA holds GPU, Omniverse, Isaac and Cosmos at the same time — what it sells is no longer just chips, but the entire underlying stack that "enables machines to learn the real world". This is the position that is least dependent on a single winner. Tier 2 · 3D Reconstruction and Digital Twin: PTC, ADSK

To turn the real world into a digital world that AI can train on, someone must first build the industrial 3D, digital twin and design workflows. They are the "data entry point" for the real world to enter the model. Tier 3 · Spatial Perception: OUST, MBLY

Lidar and machine vision — responsible for converting the physical environment into spatial data that machines can read. Without this tier, all the content above is just castles in the air; but this tier is also the most competitive, with the fiercest competition over cost and consistency. Tier 4 · End-to-End Deployment: TSLA, SYM

The players who actually deploy AI into the real world. They have the greatest upside potential, but also the most concentrated risks: a model that can run in the lab does not mean a robot that can be delivered at scale; an effective simulation does not mean it can be smoothly migrated to the real world. Going forward, do not only focus on how stunning the demo videos are, pay attention to three things: whether simulation can truly reduce field training, whether synthetic data can truly improve the success rate of models, and whether customers are willing to pay for it continuously. The last round of AI revolution turned human knowledge into tokens. The next round may turn the entire real world into computable, simulatable and trainable data.

Language allowed AI to understand humans; and 3D is what will allow AI to truly step into the real world.

This article is sourced from "US Stock Investment Network", and is authorized for release by 36Kr.