The whole world is speaking the language of AI
Smartphones have dominated the digital ecosystem for over a decade, acting as a black hole for our attention and our most intimate portable possession. But from the very beginning, phones were designed for "people staring at them" — their entire logic stops at the screen.
The demand for AI is the exact opposite: it needs to continuously perceive the physical world — see what you see, hear what you hear, be present at all times, rather than waking up only after you unlock the screen.
When AI truly becomes a fundamental capability, it will sooner or later break out of the screen and find its own form. This will be a long process of exploration and evolution.
The column "AI Artifacts Journal" was born from this idea. ifanr wants to observe with you continuously: how AI changes hardware design, reshapes human-computer interaction, and more importantly — what form will AI take to enter our daily lives?
This is the 19th article in the "AI Artifacts Journal".
In the past two years, the most obvious change in AI is not just that models have become smarter, but that their positioning has shifted from a professional tool for a few people to a fundamental capability for daily life.
Nowadays, we are even more accustomed to giving instructions to AI anytime and anywhere while walking, driving, cooking, or even slacking off at work.
Image via Google
However, as usage frequency increases, the human-computer interaction method dominated by the keyboard and mouse from the PC era is gradually revealing its shortcomings —
"Manual typing" requires users to first have an idea, then organize scattered thoughts into coherent language, and maintain a certain logical relationship in the narration, before they can turn what's in their head into prompts that AI can understand.
The Passion of Creation by Leonid Pasternak
On the other side of the keyboard, the "voice input" dominated by the smartphone era has almost no such constraints.
After all, the biggest difference between speaking and typing is that voice allows people to think, supplement, and revise while outputting, which is closer to the original way thoughts occur than keyboard input.
In addition, voice input does not require a stable desktop environment, does not require users to keep their heads down, and does not require excellent written expression skills —
For the vast majority of ordinary people who have not received systematic writing training, voice is far closer to "truly natural interaction" than text input.
Under this widespread demand, a number of products centered around the voice entry point have naturally emerged.
For example, the fast-growing Plaud focuses on packaging recording, transcription, and summarization into a complete workflow;
Typeless, which has become a model product, is responsible for removing filler words, cleaning up formatting, and "helping you complete the text".
Wearable devices like Lightwear go a step further, trying to integrate a complete full-scenario AI usage flow outside of traditional smartphones.
Image via LightSail Tech
These products have different forms, but they are all betting on the same thing:
The next-generation AI entry point does not necessarily require users to sit down, open an app, and type carefully.
Hey, AI
"Natural language interaction" sounds very advanced, but its advancement actually lies in communication efficiency, and there is nothing earth-shattering about its technical principles.
The lowest-threshold solution is a pair of TWS earbuds paired with the AI on the phone, which can form a 24/7 online virtual assistant.
The original AirPods ad "Bounce" by Apple
After all, earbuds are already in your ears, and mobile systems and third-party apps already have ready-made voice transcription, shortcuts, and cross-app input functions.
Without any dedicated hardware, you can convey your scattered commands to any AI model.
More cruelly, its speed, compatibility, and cross-device capabilities are still superior to most of the so-called "groundbreaking" dedicated AI hardware on the market —
Image via Mashable
However, the reality is that the real bottleneck of this almost zero-threshold combination is not on the phone, but in the microphone's processing of ambient human voices.
Take AirPods as an example. Due to AirPods' microphone strategy that emphasizes natural sound quality and is reluctant to apply aggressive human voice noise reduction, in multi-person environments, it often picks up surrounding voices and transcribes them into text as well.
This is not a problem for making phone calls or sending voice messages, but for AI input, once the prompt is "mixed", it is easy to get completely irrelevant responses:
In contrast, TWS earbuds from domestic brands generally adopt a more aggressive human voice isolation strategy, sacrificing some transparency mode naturalness in exchange for better accuracy during phone calls and AI voice input.
If we look deeper, this means that at a time when AI has become a "mainstream usage scenario", the evaluation criteria for good and bad earbuds are changing —
In the past, earbuds competed on sound quality, latency, and noise reduction. In the future, in addition to these, they will also compete on the intensity, accuracy, and separation capability of human voice noise reduction.
We can even say that earbud microphones are gradually transforming from a long-forgotten call accessory into the most front-end AI controller.
Mission: Impossible (1996)
The second route for voice control is smart glasses.
Although smart glasses always emphasize "adding a HUD to your life", from a commercialization perspective, the only successful smart glasses are still the basic models with voice, photography, and open audio as their core features.
Image via Laptop Mag
The reason behind this is not complicated: the glasses display system sacrifices weight, battery life, and cost, and requires cultivating a whole new set of user habits from scratch.
And "ask whatever you see, and AI will tell you" is itself enough to form a relatively clear selling point for AI.
Image via Meta
But the problem with smart glasses is that their "unobtrusiveness" mostly exists in promotional videos.
After all, when a pair of glasses with a camera is continuously facing others, it is difficult for the other party to determine whether the device is recording or picking up audio, and this uncertainty itself creates social pressure.
In addition, limited by size and battery life, the microphone array, call noise reduction, and speaker performance of smart glasses may not necessarily be better than mature TWS earbuds.
Image via Stuff
More importantly, in the current sales model of smart glasses, these products are often deeply tied to the brand's client and cloud models, leaving users with no freedom in model selection, data migration, and cross-device switching.
The simplest example is Meta.
Meta Ray-Ban smart glasses are well-balanced in all aspects, but they force users to connect to Meta AI and cannot be used as Bluetooth audio devices to directly wake up Siri or other smart assistants.
Image via Engadget
This level of binding makes smart glasses a representative product of "voice AI interaction", but it is difficult for them to become the mainstream product.
The third route is the one that has gradually emerged since 2026 — creating an independent hardware entry point for voice input AI.
For example, the recently released Codex Micro, a collaboration between OpenAI and Work Louder, is an external console for AI agents.
It is specifically equipped with a push-to-talk (PTT) button for voice input:
Image via OpenAI
Although the Codex Micro itself does not have a microphone, and the PTT function needs to call other audio input devices connected to the computer, this design is still symbolic:
Voice input is no longer just a small icon in the UI, but has gradually begun to occupy a stable physical position similar to the right mouse button, the phone's camera button, or the keyboard's fingerprint sensor.
Setting voice input as a separate button basically acknowledges the high-frequency nature of this AI interaction method — for some users, it may even be used more frequently than the 26 letter keys.
Luo Yonghao demonstrating TNT
The further vision of PTT/TNT is to completely get rid of the screen.
ifanr previously broke the news that among the several hardware products OpenAI is planning, there is a screen-less portable device that understands the user's context through voice, camera, and environmental sensors.