Why does a 3D printer hide the greatest ambition of GPT-6?
The promotional video for the release of GPT-6 Astra starts with a yellow circle:
The person in the footage asks Astra to draw a yellow circle, then turn it into the porthole of a rocket. A few seconds later, the outline of a rocket grows around the circle:
He then asks GPT to add more details, and turn the entire rocket into a 3D model in Blender:
But the task is not yet finished. He returns to the rocket model again and gives an instruction: Now generate a file that I can send to a 3D printer.
So GPT-6 generates an STL file and sends it to the Bambu Lab P1S on the desktop — even though the footage is pixelated, it is still instantly recognizable — a 3D printer released in the same year as GPT-4.
The rocket on the screen is turned into a physical entity in the real world through this P1S printer.
From the yellow circle, to the rocket sketch, to the 3D model, and finally to the physical object in reality — in a short film of less than three minutes, OpenAI delivers a perfectly circular, well-coordinated narrative.
If the only goal was to prove that GPT-6 can do 3D modeling, being able to use Blender would already be enough. But OpenAI deliberately ends the story with a 3D printer.
Because this time, AI is no longer just generating content on the screen — it is starting to step into the physical world.
GUI, Possibly the First Lesson for AI to Understand the World
GUI, or Graphical User Interface, is a concept that has existed for more than half a century. We click on windows and drag icons every day, rarely stopping to think about what it really is.
But from the underlying logic, GUI has always been doing one thing: translating the complex, invisible logic inside computers into a world that humans can understand and operate —
You see a "desktop" as soon as you turn on your computer, files are stored in "folders", apps are displayed in "windows", and unwanted items go into the "trash".
In this system, every object has attributes, and there are relationships between different objects. Humans can perform a wide range of actions to bring the entire system into a new state.
Essentially, GUI is a highly abstract representation of the real world.
This is exactly why the improvement of GPT-6's GUI understanding capability is far more significant than "AI being able to click buttons for me".
Previous vision models that saw Photoshop only needed to tell users where the lasso tool is; but now that AI is set to truly complete Computer Use, it must understand what file it is currently processing, which object needs to be modified, what results the operation will bring, and where to roll back if something goes wrong.
This is a complete "Perception → Judgment → Action → Feedback" loop.
This capability has a subtle isomorphism with the real world —
The real world is also made up of objects: a cup has a position and weight, a door can be open or closed, and turning the steering wheel will change the angle of the wheels. Humans observe objects in the environment, judge their relationships, perform actions, and then adjust their behaviors based on the results.
AI in the past learned "how the world is described", while modern AI has started to learn "how the world is operated".
The gap between the two is action.
Intelligence can only be turned into practical capability when it is put into an environment. And GUI is exactly the largest, most standardized "environment" ever created by humans.
From Operating Software to Operating the Real World
Looking back at the P1S printer in the promotional video, its meaning has become completely different.
3D printers happen to sit on the boundary between the GUI world and the real world.
Generated images are still pixels, generated text is still characters, and most running code stays inside the computer.
But when an STL file is imported into a 3D printer, a completely different transformation takes place — the "object" in the digital world becomes a "physical entity" in the real world.
In this sense, the 3D printer is the portal that breaks the boundary between the digital world and the physical world.
Looking back at the technological development over the past few years, you will find an interesting parallel.
In 2022, Bambu Lab launched its first-generation X1 series, and at the end of the same year, ChatGPT built on GPT-3.5 went online.
The two paths started separately, but both have been shortening the distance between ordinary people and complex technologies.
Four years later, when the model generated by GPT-6 is printed into a physical object by Bambu Lab, the two paths converge on this tiny rocket.
What GPT-6 eliminates is the complexity on the software side — I don't have to follow tutorials step by step to learn Blender or CAD. I only need to tell the model what I want to get in the end, and the model will deliver it to me.
What Bambu Lab eliminates is the complexity on the manufacturing side — I no longer need to understand the hot bed temperature and infill density, I only need to load the consumables, import the drawing, and press the print button.
The former frees people from learning complex software, the latter frees people from learning complex machines. Strip away the unnecessary parts, and what remains is creation itself.
A natural language request, passing through GUI, design software and manufacturing equipment, finally turns into a real object — this is the complete path from intent to reality that OpenAI demonstrates in its promotional video.
Over the past decades, human-computer interaction has been a field that keeps pace with the times, whose research goal is to make the digital world closer to the real world:
Command lines require humans to learn the language of machines, GUI translates machines into the visual world familiar to humans with icons and windows, and touch screens turn "moving the mouse" into directly tapping on objects...
Large language models, however, have begun to subvert this discipline of human-computer interaction — we no longer have to find buttons and open menus by ourselves. AI can read GUI and control devices such as 3D printers to complete operations for humans.
Software will not disappear, but people may no longer need to rack their brains to learn it — more and more people will no longer need to directly work with Photoshop, Blender, CAD...
Just like people using cameras today do not need to understand how the ISP processes each pixel, people using CAD in the future will not necessarily need to know how many menus and tools are involved in a modeling process.
What we really want is never a CAD file, but a bracket, a chair, a part, or a tiny rocket that we can hold in our hands.
GUI is the language that humans use to translate the real world for computers, and now AI has fully understood it.
And 3D printers translate the abstract concepts inside computers back into reality.
From GPT to the 3D printer, that is the distance between a single thought and tangible reality.
This article is from the WeChat Official Account ifanr (ID: ifanr), written by Xiao Qinpeng, and republished by 36Kr with authorization.