"The US version of the Doubao Phone" is totally lame.
On the morning of August 13 Beijing time, Google's annual hardware event Made by Google 2026 was held as scheduled. The event unveiled a full lineup of new hardware including the Pixel 11 series smartphones, Pixel Watch 5, and the brand-new Pixel Tag tracker, with a preview of its upcoming smart glasses set to launch before the end of the year.
However, after watching the nearly two-hour event, many audiences described the overall experience as "underwhelming".
To start with, the hardware upgrades are very incremental. The most notable new feature of the new products at this launch event is the small circle of LED lights added to the back of the Pixel 11 Pro series, which Google calls Hilight. Its core function is to support color-coded notifications for incoming calls from specific contacts, with no support for SMS alerts or third-party app notifications.
Software improvements are relatively more substantial. Gemini can independently complete search operations across in-phone apps such as information services and maps, with optimized details of input functions and camera performance. As a result, the public widely notes that the Pixel 11 is not a phone designed to win users over by hardware specs, but rather the "physical vessel" for Gemini. This concept sounds highly similar to China's "Doubao Phone", which is built entirely to deliver AI-powered services.
This has also sparked widespread rethinking across the industry: when Google turns all its hardware into containers for Gemini, is it still selling smartphones, or an AI assistant that users can call on at any time? How many features of this "US version of Doubao Phone" are original innovations from Google, and how many are following the footsteps of its competitors?
01. Google's Hardware Matrix: Filling in the Gaps, But No Standout Breakthroughs
At this year's Made by Google event, Google showcased a more diverse product portfolio than previous years. The headliner is the full four-model Pixel 11 smartphone lineup, alongside the Pixel Watch 5, the new Pixel Tag tracker, and a teaser for smart glasses to be launched within the year. Despite the rich variety of products, none of them can be described as stunningly innovative.
First, let's look at the most anticipated Pixel 11 series smartphones.
In terms of pricing, all four new models are $100 more expensive than their predecessors. However, since the 128GB storage variant is removed for both Pixel 11 and Pixel 11 Pro, the price of their 256GB versions remains the same as the previous generation, while the Pixel 11 Pro XL and Pixel 11 Pro Fold see a $100 price hike respectively. Judging from the overall specs, the hardware of the Pixel 11 series is mostly minor annual updates. The entire lineup is equipped with Google's new generation Tensor G6 processor, which is custom-built for on-device Gemini operation. The Pro series is upgraded with a new telephoto lens supporting up to 120x zoom, while the standard model retains the exact same camera specs as its predecessor. For battery performance, the standard model, Pro, and Fold support 30W charging, and the Pro XL supports 45W charging, making the overall battery performance fairly average.
The biggest hardware highlight is the new LED indicator on the back of the Pro series, officially named Hilight. Its core functions are twofold: on one hand, it shows the running status of Gemini, for example, it stays solidly lit when the user is interacting with the AI assistant. On the other hand, it displays the user's custom assigned color for incoming calls from specific contacts. No other functions are available for now: it does not support alerts for SMS or other third-party apps, nor does it reserve open interfaces for developers.
Source of the Hilight feature / Google official website
Marques Brownlee, a tech reviewer who got the hands-on unit in advance, commented on this series with the phrase "it is what it is" and used this line as the title of his review video, which shows that the public is not fully satisfied with this new feature.
Next up is the Pixel Watch 5, priced starting at $399.
It runs on the overclocked version of Qualcomm Snapdragon W5 Gen 2, with a 12% faster CPU performance. But the real highlight is not the CPU: the on-device RAM is upgraded from 2GB to 3GB, a full 50% increase. This brings very tangible improvements: the Gemini capabilities on the watch are greatly enhanced, for instance, users can create custom watch faces, and activate Gemini to have a conversation directly when they raise their wrist.
Last, let's look at the Pixel Tag and the upcoming smart glasses.
Pixel Tag is Google's first entry into the item tracker market, priced at $29 for a single unit and $99 for a pack of four, directly matching the pricing of Apple AirTag, and it will go on sale in November. There is not much to note about the product itself, but its launch fills a gap in Google's hardware ecosystem: from smartphones to watches, earbuds, and item trackers, Google now has a complete hardware lineup that can rival Apple's ecosystem. The "smart glasses" is the hook Google threw out at this event, no physical product was shown, only a preview that it will be launched within this year.
It is clear that Google presented a full matrix of AI containers at this launch, but none of the individual products bring major surprises. The price hike of Pixel 11 is backed by doubled base storage, but its new features are not impressive enough; the RAM upgrade of Pixel Watch 5 is a very practical improvement, but its functions do not break away from the existing framework of other smart watches; Pixel Tag is a typical market follower product, and the smart glasses are not even unveiled yet.
02. How is Google's AI Smartphone Performing?
The real protagonist of the entire launch event is Gemini. It has evolved from a standalone assistant to be fully integrated into the system service layer of Pixel, taking over everything from information lookup to cross-app task execution. With Gemini at its core, what level of capabilities has Google's AI smartphone reached? This event has given the answer.
In terms of "task execution" capabilities, Gemini has gone far beyond simple chatbot interactions.
In the official demo, it can read information across apps including SMS and calendar, automatically search on Google Maps as requested, find suitable restaurants, place orders and make reservations, and even make direct phone calls to the merchants. It can also proactively suggest flight information, restaurant lists, and remind users of calendar events.
Image source / Google official website
However, all these smooth experiences are currently limited to Google's own services, and to ensure user security, key steps still require user confirmation. This function does not differ much from the AI capabilities of other AI smartphones on the market.
Voice dictation is one of the biggest highlights of this event.
The new Rambler feature on Gboard, powered by Gemini, is officially launched with three core functions: automatic text polishing, intent correction, and one-click reversion. It can automatically filter out filler words like "um" and "ah", automatically resolve logical contradictions even if the user changes their mind mid-speech, and undo previous content when the user asks it to, which is quite appealing to business professionals and student groups.
In addition, it is also equipped with features including sign language recognition, real-time translation, and imaging AI.
At the event, deaf actor Daniel Durant demonstrated real-time sign language to text conversion: the front camera of the phone captures his sign language movements, the large language model translates the sign language into text, and the result is directly displayed in the real-time transcription. The real-time translation function is now integrated into YouTube and podcasts, allowing users to dub the audio content into another language in real time when watching videos or listening to podcasts.
In terms of imaging, after the user presses the recording button, Magic Capture will analyze hundreds of frames before and after the moment, and pick out the best several shots. The new Creator Suite added to the camera also comes with a built-in teleprompter: users can input their script in advance, and the text will automatically scroll along with the user's speaking speed and natural pauses after they hit the record button.
All these features are powered by the combination of Gemini's multimodal model and the on-device computing power of Tensor G6. Although Google did not elaborate on the underlying tech stack during the event, it is confirmed that features including sign language recognition, real-time translation and video understanding all rely on strong on-device AI capabilities.
But all these features share one thing in common: most of them are "catching up" innovations rather than leading breakthroughs. For example, in the field of sign language recognition, Apple has already supported sign language detection in FaceTime. In the field of real-time translation, Apple and Microsoft mainly cover text and voice scenarios, while Google extends real-time dubbing to specific content formats including YouTube videos and podcasts.
03. Is It the US Version of Doubao Phone?
After the launch event, the claim that "the Pixel 11 is the US version of Doubao Phone" has been widely circulated. This metaphor captures some similarities between the two products, while missing many key differences.
From the strategic perspective, the Pixel 11 and Doubao Phone are indeed following the same development path. Facing the pressure from competitors like OpenAI in model benchmarking, Google stated at its I/O conference that it has shifted its focus to AI agents. Gemini is deeply integrated with Google's full suite of services, while Doubao is deeply bound to ByteDance's product ecosystem. Neither of the two parties is chasing the peak performance of a single technology, but focusing on expanding the presence of their AI in more user scenarios.
However, the technical routes of the two products are not the same.
A practitioner in the large language model industry introduced that the AI function that allows the AI to operate your phone on your behalf can be divided into two levels.
The first level is the "operation layer", which has two different approaches: the first approach is that the AI "operates by itself", monitoring the screen and clicking step by step following the user's usual operations, such as opening apps, locating buttons, and tapping confirm, which is called GUI simulation tapping. The second approach is that the AI does not interact with the screen, but directly sends instructions through the pre-reserved internal interfaces of the apps, which is called API calling.
The second level is the "collaboration layer", also known as A2A, which allows the agent inside the smartphone to communicate with the system agent, and hand over the task to the system agent for completion. This is a more advanced architecture, but there is no mature large-scale implementation in the industry for now.
Specifically for these two products, both Pixel and Doubao Phone belong to the operation layer. The Pixel prioritizes first-party API calling, and falls back to GUI simulation tapping as a backup for apps that do not open their interfaces. The first generation of Doubao Phone adopted the pure GUI route, but due to risk control restrictions from top apps and public concerns over security, the second generation of Doubao Phone has shifted to the A2A architecture.
Image source / pixabay
The difference between the two routes does not lie in which one is more advanced, but in the different depth of vertical integration.
The Pixel 11 uses Google self-developed hardware, custom Tensor G6 chip, and runs self-developed Gemini model. The full stack from the chip to the operating system and the AI model is fully controlled by Google. This allows Google to make optimizations at the underlying level, such as adding RAM to the watch to run AI models smoothly, and ensuring the on-device AI runs smoothly even without internet connection. This kind of vertical integration of "rebuilding hardware for AI" is very difficult for Doubao Phone, which relies on third-party hardware manufacturers, to replicate.
Doubao Phone currently mainly cooperates with hardware vendors such as Nubia, and has limited say in the hardware design process. Its advantage lies in fast software iteration and deep scenario adaptation, but the upper limit of its on-device optimization is restricted by the hardware capabilities of its partner manufacturers. However, in terms of engineering implementation, Doubao is indeed a pioneer in the field of GUI automation.
In short, the metaphor of "the US version of Doubao Phone" is valid at the strategic level, but not accurate enough at the product level. A more precise description is that they are two different solutions under the same industry trend, with consistent core goals but different trade-offs in implementation methods.
Back to the question raised at the beginning: is Google selling smartphones or AI? The answer is that Google may no longer intend to draw a clear line between the two. The Pixel series is gradually becoming the "physical vessel" for Gemini. Hardware is only a one-time entry point, and the continuous subscription services and capabilities powered by AI are the longer-term business.
This article is from WeChat official account "AIX Finance", written by WANG Lu, edited by WEI Jia, published with authorization from 36Kr.