OpenAI drops a bombshell: its function that can generate a full website with just one single sentence is completely revolutionizing and disrupting the entire SaaS industry.
Just a moment ago, OpenAI has taken another evolutionary leap.
Right now, the highly anticipated ChatGPT Sites is officially open for public beta!
Even if you do not know a single line of code, you only need a few natural language prompts or upload a sketch, and you can generate a highly professional, feature-rich interactive website in just over ten minutes as easily as writing a Google Doc.
Front-end development can now be completed using nothing but natural language, and the traditional website-building SaaS industry is facing massive disruption.
Moreover, another new piece of news has been revealed.
On the underlying technology side, GPT-6 has demonstrated astonishing cross-modal capabilities.
With no audio playback and a completely disconnected network, relying solely on a Mel spectrogram (a visual waveform representation of sound), GPT-6 can identify the call of a deep-sea blue whale, and even recognize the lightsaber sound effect from Star Wars.
This indicates that AI's multi-modal understanding has completely broken through human sensory barriers.
It no longer clumsily "imitates" human seeing and hearing, but has evolved the ability of "native decoding" of the data structure of the physical world!
ChatGPT Sites Public Beta: Building Websites Is Now As Easy As Writing Documents
Once upon a time, owning a first-class personal website was a privilege reserved for programmers and professional designers.
Later, Dreamweaver emerged, and then drag-and-drop website-building SaaS tools such as Wix, Squarespace, and even Tilda came onto the scene.
But today, the public beta launch of ChatGPT Sites has completely shattered the barriers to website building.
After actual testing, well-known tech media Fast Company and Jeremy Caplan, host of Wonder Tools, stated that the experience brought by ChatGPT Sites is extremely disruptive!
"Now, creating a beautiful website is as simple as making a Google Doc."
At present, this feature is fully available to ChatGPT Plus, Pro, Business, Enterprise and Edu users.
How does it work? The process is incredibly simple.
1. Describe your vision: Write down the type of website you want and its target audience in a few sentences.
2. Feed it materials: If you have a screenshot of a webpage you like or a hand-drawn sketch, upload it directly.
3. Put forward functional requirements: Want link buttons? Embedded visual effects? Or complex interactive elements? Just tell it directly.
4. Have a coffee and wait: Depending on the level of complexity, a complete website will be generated in about 10 to 15 minutes.
Note that this is no longer simple "template filling", but real Vibe Coding!
Early testers have already used it to build a number of amazing projects.
First up are practical pages.
Some people have used it to create a perfect media kit sponsorship page, a modern-designed San Francisco event guide, and even a modern online picnic supply store that supports dragging products into the shopping basket.
In addition, artist Drue Kataoka used it to generate an art creation tool called GoalFlow.
After users select a national flag, they can "paint" with the colors of the flag on a fluid simulation canvas, where the paint will naturally swirl and drift like pigments in water.
This kind of front-end page involving a complex fluid physics engine used to take senior front-end engineers weeks to complete, but now it only takes a few prompt words.
Finally, there are physics engine games.
You can even create a Jenga-like game called "Glass Towers" with realistic physical effects in a few words; or make an arcade mini-game "Paper Glider" that lets you control a paper plane to fly through hoops.
Even more impressive is its iterative capability.
After ChatGPT generates the first draft, you can directly use the "annotation tool" in the sidebar to circle parts of the webpage just like grading assignments: "Change the color of this button", "Replace the font here with a sans-serif font", "Replace this image".
It will not only understand your request, but also make seamless modifications.
Clearly, industry reshuffling is now inevitable.
While early AI website building tools (such as Lovable, Bolt) still have their flexibility, and Claude's Artifacts and Claude Design also perform excellently in interactive UIs, the trump card of ChatGPT Sites is that it is integrated into an extremely large ecosystem.
If you are working in a "ChatGPT Project", it will automatically read your brand style guide and logo library, so the generated website naturally carries your brand DNA.
Right now, front-end engineers have been demoted to "prompt product managers", and those traditional SaaS companies that still rely on "drag-and-drop website building" to charge high monthly fees will probably not survive this winter if they cannot transform quickly!
Dimensionality Reduction on the Application Side: Eliminate All "Operations" and Keep Only "Intent"
Looking back, from the command line (CLI) in the PC era, to the graphical user interface (GUI) in the Windows era, to the natural language interface (LUI) in the current large model era, the history of human-computer interaction is a history of continuously lowering the operation threshold.
The essence of ChatGPT Sites is to reduce front-end development from "engineering" to "expression".
In the past, when you wanted to build a website, your brain needed to go through multiple layers of translation.
I want to sell products (intent) -> I need a shopping cart and payment system (product logic) -> I need to write HTML/CSS/JS (front-end code) -> I need to connect to the database (back-end logic)
Now, ChatGPT Sites eliminates all the hassle in the middle.
The user input end is completely reduced to pure natural language, which is the most original "intent".
The core barrier of traditional software SaaS is "we encapsulate the code into beautiful buttons for you to click", while the dimensionality reduction strike of large models is: you only need to speak your mind.
From now on, all transitional tools centered on "making it easier for non-programmers to write code/build websites" will lose their living space in this "end-to-end" generation era where intent directly generates results.
GPT-6 Recognizes Spectrograms At A Glance, Breaking Human Sensory Barriers
Moreover, today, a set of tests released by AI researchers ChrisGPT and Max Rubin has shocked countless audio engineers and AI practitioners.
This test is very simple: convert sound into an image, and let GPT-6 read the image directly.
ChrisGPT uploaded a Mel spectrogram to GPT-6, a technology that visualizes the frequency, time and energy of sound into a two-dimensional image.
During the process, he did not play any actual audio, and the large model was in an offline state.
The only prompt was "This sound comes from an animal/mammal in nature."
GPT-6 replied: "Whale, probably a blue whale."
The entire internet is stunned.
Because it not only answered correctly, but also exceeded our cognition — the reasoning logic behind it has gone far beyond simple image matching!
It should be noted that blue whales produce extremely low-frequency calls.
But in nature, the calls of lions, tigers and elephants also fall in the low-frequency range of 40-200 Hz. Relying solely on a mass of energy near 50 Hz on the spectrogram, it is impossible to conclude that it is a whale.
How did GPT-6 do it? Audio engineers analyze that GPT-6 keenly captured the "morphological characteristics of energy that persists over time".
It is not looking at a single point, but understanding how sound waves flow and decay over time in this specific physical space.
With very little context, zero-shot guessing the animal category directly from a "picture of sound" is an incredibly impressive capability.
The demonstration by another tester, Max Rubin, is even more memorable.
He tested non-natural sound effects, and the results showed that Astra could directly zero-shot identify the sound of a lightsaber being swung from Star Wars from the Mel spectrogram!
Max exclaimed: "We haven't even scratched the surface yet."
Although in the field of scientific research, the academic community has long tried to combine spectrograms with visual models to detect whale calls, and even train dedicated LLMs to read Mel graphs for speech tasks.
However, this is the first time a general-purpose large model has been publicly confirmed to be able to perform this level of cross-modal reasoning without any special fine-tuning.
This is like, you have never taught a person how to read sheet music, you just show them a full orchestral score.
As a result, they can not only "hear" the sound in their mind, but also accurately tell you that the trombone in the second measure played a wrong note.
Large Models Are Breaking Free From The Constraints Of Human Biology
The human perception of the world has huge limitations.
We must rely on air vibrations to vibrate our eardrums before the brain can convert the physical vibrations into electrical signals of "sound". In the human subconscious, sound is something that is "heard" with the ears, and images are something that is "seen" with the eyes. This is the "sensory barrier" set for us by millions of years of evolution.
Early multi-modal AI was actually "imitating" humans. We let it listen to audio files, show it images, and try to make it understand like a human.
But the significance of GPT-6 Astra identifying spectrograms is: large models are breaking free from the constraints of human biological organs!
For GPT-6, the long call of a deep-sea blue whale and the tearing sound of a Jedi swinging a lightsaber are essentially no different — they are both just forms of energy fluctuation in the physical world.
When this fluctuation is converted into a mathematical and geometric representation such as a Mel spectrogram, GPT-6 can directly "read" the physical rules themselves.
It does not need to grow ears, it does not need to "hear" the sound, it can directly "see" the physical formula of the sound.
This means that AI's multi-modal understanding has evolved from "imitating human senses" to "native decoding" of the data structure of the physical world.
Today it can look at a spectrogram and recognize a blue whale; tomorrow, if you show it an image of a seismic wave, can it directly predict the next fault rupture? If you show it an image of brainwaves, can it directly read what you are thinking?
Humans perceive the projection of the universe through their physical bodies, but now AI is directly reading the source code of the universe.
References:
https://x.com/ChrisGPT/status/2097047569520115800
https://x.com/maxxrubin_/status/2096892510241268094
https://www.fastcompany.com/91602695/chatgpt-sites-makes-building-a-website-feel-like-making-a-google-doc
This article is from the WeChat official account "AI Era", written by ASI Revelation, edited by Aeneas David, and republished with authorization from 36Kr.