Why is OpenAI's first AI hardware a screenless "donut"?
OpenAI's first AI hardware is arguably the most mysterious Schrödinger's product of the past two years. We still have no idea what it actually looks like, but a new update on it pops up every few months.
Early rumors claimed it was a smart pen, then desk lamps, glasses, and wearables were all named as the possible product in turn. As the rumors piled up, even we onlookers developed a sort of "cry wolf" PTSD.
According to Bloomberg's latest report today, the first consumer device that OpenAI is advancing in collaboration with Jony Ive's team is currently closer to a screenless smart speaker, shaped as a hockey-puck-sized ring.
Rendering produced by @PatentlyApple
To put it more plainly, it looks like a metal donut.
If you only look at the product form, this product that may be priced at 300 to 400 US dollars can easily be mistaken for a more expensive, more exquisite smart speaker. But OpenAI has a more radical positioning, hoping to make it a kind of AI-first computer, that is, a computing device that takes AI as the primary interaction entry.
The hockey-puck-sized disc is only the first form that has surfaced so far. What OpenAI really wants to verify is: when AI is capable enough to undertake most interactions, do we still need to stare at a screen every day?
A moving metal "donut"
The information revealed so far is quite specific.
The whole device is ring-shaped, close to the size of a hockey puck, can be picked up with one hand and moved between different spaces in the home. It has no display screen, but is equipped with speakers, microphones, lights, cameras and other environmental sensors, and is powered by a battery, expected to be released next year.
Users can place it on the bedside table, or bring it to the kitchen, living room, or even hold it directly in their hand for use.
It is worth noting that this form also highly matches the suspected leaked device of the well-known American designer Joe Gebbia that was photographed earlier.
More interestingly, it is not a completely static piece of metal.
The report states that a movable mechanical structure will also be designed inside the device. When the user talks to the AI, some components will move, accompanied by light changes, to let people know that it is listening, responding or performing tasks.
Smart speakers have long had a problem. Even if Alexa, Google Assistant and Siri can sustain conversations, users are still facing a static electronic product. You call it, it lights up; you ask it a question, it falls silent again after answering.
OpenAI hopes to make the device have a stronger sense of presence through movements, lights and more natural voice interaction, and even make users feel that it is like an object that truly participates in the surrounding environment.
And the cameras and sensors also mean that the information it obtains will not be limited to sound. Traditional smart speakers can only know what you said, while OpenAI hopes it can further understand what is happening around you.
Rendering produced by Trung Phan
For example, you may no longer need to tell the AI that there is a bottle of medicine on the table, but directly ask when to take this bottle of medicine; you do not need to describe the situation in the kitchen in detail, but directly ask if the food in the pot should be turned off the heat.
If vision, voice, long-term memory and environmental perception are truly combined, what AI gets is no longer just a single conversation, but a continuous existing context.
This is also the direction that OpenAI has been pushing for the past few years.
The truly valuable AI in the future will most likely not only answer questions, but also know what the user is doing and what they have done in the past for a long time, and continue to provide assistance accordingly.
After leaving Apple, Jony Ive is searching for the next product after iPhone
Another thing that makes the whole incident more dramatic is Jony Ive.
The design partner for OpenAI's first hardware is none other than his LoveFrom studio.
As the person who once led the design of iPhone, iPod, iPad and Apple Watch, Jony Ive has almost participated in defining the most important product language of consumer electronics in the past two decades.
Now, he is working with OpenAI to find another possibility.
Last year, OpenAI acquired io, the hardware company co-founded by Ive, and the two sides formally moved from cooperation to deep binding. OpenAI has also publicly expressed its plan to build a complete device business in the long run, and the ultimate goal even includes creating a product that can replace smartphones.
At the same time, the relationship between Apple and OpenAI has deteriorated rapidly.
Recently, Apple has sued OpenAI and relevant personnel over trade secret issues, accusing the other party of obtaining and using information involving product design, manufacturing processes and supply chains. One of the points of contention is whether OpenAI ever asked suppliers that have long-term cooperation with Apple to develop metal surface treatment technologies similar to those used by Apple.
OpenAI denies the relevant allegations, considers Apple's claims too broad, and later emphasized in the document requesting the court to dismiss the lawsuit that it is building something completely different from Apple's existing products.
At least judging from the product form exposed so far, this ring device is indeed difficult to directly correspond to any existing product of Apple. The routes of the two companies have shown very obvious differences.
Apple's new generation of home devices under development, as reported, is closer to a display screen connected to a mechanical structure, which can turn to the person speaking and undertake visual tasks such as FaceTime.
Apple is still trying to bring display screens into more scenarios, while OpenAI is trying to verify a more radical question: if AI is powerful enough, does the human-computer interaction in the future still require the screen to exist all the time?
In a sense, this is also a very interesting turning point in Jony Ive's career.
Nearly 20 years ago, Jony Ive participated in the design of a touch screen that almost swallowed the entire consumer electronics industry; 20 years later, he is participating in the design of a computing device that tries to remove the screen.
Why OpenAI's first hardware happens to have no screen
The core value of smartphones in the past decade or more is largely built on screens and apps.
Users search for information through the screen, call services through apps, and then complete each operation step through touch. The entire business model, product design and even the advertising system of the mobile Internet are all built around this interaction method.
Agents are trying to change the most fundamental layer of the relationship. When users only need to express their intentions, and AI can independently decide what software to call, what information to read, and which operations to complete, the interface value of apps will decline, and the tasks undertaken by the screen will also decrease accordingly.
Rendering produced by X blogger Jesse Waites (translated into Chinese)
The so-called AI-first computer of OpenAI is probably built on this judgment.
Future computing devices may no longer focus on waiting for users to click, but gradually become agents that can listen, see, remember context, and actively help users complete tasks.
From this perspective, having no screen has instead become a very clear product expression.
Pulling out your phone, unlocking it, opening ChatGPT, and entering instructions are still interaction methods belonging to the smartphone era. What OpenAI wants to do is to further eliminate these startup steps, and let AI exist in the environment for a long time.
When you think of something in the kitchen, you can speak directly; when you suddenly remember tomorrow's schedule before going to bed, you don't need to touch your phone; after the device obtains environmental information through cameras and microphones, it can even save a lot of background descriptions.
In contrast, smart glasses are closer to all-weather AI, but privacy issues are also more sensitive.
A continuously worn camera will bring all the people around into the sensing range. The attempts of smart glasses in the past few years have proved that technical feasibility does not equal social acceptability.
In contrast, devices placed in home spaces are easier to be accepted.
OpenAI's entry from the smart speaker form is probably to avoid the most sensitive privacy disputes of wearable devices while testing all-weather interaction.
Of course, AI hardware in the past few years has left enough negative examples. The most typical problem of Humane Ai Pin is that removing the screen itself will not create the next generation of computing platforms.
When model capabilities, response speed, agent execution capabilities and product reliability cannot match, having no screen will even amplify all interaction problems.
The difference between OpenAI and these startups is that it already has models and users, and now it needs to find a physical carrier more suitable for these capabilities to exist, which also reflects the difference between AI+ and AI-Native.
The logic of AI+ is to "add a layer of intelligence" on existing devices and existing interaction methods. In essence, people actively open the tool, and the tool passively responds to demands.
The idea of AI-Native is completely opposite. It does not patch on the existing system, but redefines "how humans interact with computing". In this framework, the model is the interaction itself; the device is no longer an entry, but part of the environment.
When this paradigm is established, hardware is no longer just a container for carrying software, but rises to become the primary constraint that determines the form of interaction. Once the technology platform changes, what is most worth watching is often that the interaction entry begins to shift.
Almost every platform change in history is accompanied by the migration of entries: the entry in the PC era is the desktop and browser, the entry in the mobile Internet era is the app icon and touch screen, and the entry in the AI era is changing from clicks to conversations, from operating systems to intent systems.
Accompanied by this, the distribution logic, product form and even business model of the original ecosystem will be reshuffled. The cost of breaking the situation is huge, but the fruits that can be picked after breaking the situation are also the most lucrative.
Many truly important products in the history of technology originally came from a seemingly redundant question: Why must a computer be placed on a desk? Why must a phone be connected to a wire? Why must the Internet be accessed through a browser? Why must a mobile phone have a screen?
Looking back many years later, it is often these unnecessary questions at the time that push technology forward. For things that are already good enough, there are always people wondering if there is another possibility? Because even without OpenAI, someone else would definitely push open this unknown door.
In a sense, human civilization also started from this kind of curiosity.
Hundreds of thousands of years ago, when a primitive man put down what he was doing for the first time and looked up at the stars that seemed to have nothing to do with eating and survival, he certainly did not know the answer, nor did he know where this glance would eventually bring humanity. But everything that happened afterwards probably started from that glance up.
This article