Android has entered the era of 2nm intelligent agent chips overnight, and Qualcomm has released the world's fastest mobile CPU.
Mobile phones are at another turning point of transformation. The keyword this time is Agent.
Based on this judgment, Qualcomm has just unveiled two brand new flagship chips —
Qualcomm Snapdragon 8 Gen 6 Ultra and Qualcomm Snapdragon 8 Gen 6 Super Ultra.
Not only the CPU, GPU and NPU are fully upgraded around agent scheduling and on-device large model deployment.
There are also many large model vendors among its partners.
For example, QwenBook, the Qwen tablet that makes its debut overseas for the first time.
And the Doubao phone.
It has also cooperated with Step, Wulianhuo, Longsys and other partners to realize complete on-device deployment of the 30-billion-parameter MoE model.
Without further ado, let's take a look at how Qualcomm this time "makes intelligence closer to the processor".
Dual Flagship Chips Targeted At Agents
Both the Snapdragon 8 Gen 6 Ultra and the Super Ultra version are built on the 2nm process technology.
The Snapdragon 8 Gen 6 Super Ultra is positioned as Qualcomm's "most powerful mobile platform", equipped with "the world's fastest mobile CPU".
The Ultra version aims to bring the same 2nm, Oryon CPU, Adreno GPU and Hexagon NPU architecture to a wider range of high-end flagship devices.
The World's Fastest Mobile CPU
The maximum main frequency of the new generation Qualcomm Oryon CPU reaches 5GHz.
Qualcomm claims that with the combined effect of 5GHz and improved IPC, the new Oryon CPU is "the world's fastest mobile CPU". Compared with the previous generation, its performance is increased by 13%, energy efficiency by 37%, and response speed by 15%.
What deserves more attention on the CPU side is Oryon FlexCache.
Qualcomm emphasizes that this brand new cache architecture has expanded the way Qualcomm chips build memory hierarchies.
Essentially, it reworks the allocation of cache resources: in the past, the relationship between cache capacity and cores was relatively fixed, while FlexCache turns the cache into a dynamically allocated cache pool. When a super core suddenly undertakes a high-load task, it can use "almost the entire cache pool" on demand while still maintaining low latency.
For agents, this means that the response speed can still be guaranteed as much as possible when performing complex tasks with frequent scheduling and suddenly increased load.
GPU: Matrix Cores Built Exclusively for AI
The performance of the new generation Adreno GPU is increased by 44%, and energy efficiency by 40%. Qualcomm says this is "the biggest generational leap ever" for Adreno GPUs.
This time, Qualcomm has added Matrix Cores dedicated to AI workloads inside the Adreno GPU, so that neural network computing can be completed directly within the graphics rendering pipeline, without having to move the data out for processing by other computing units.
With Matrix Cores, Qualcomm has further launched a very important new feature: Adreno Neural Fusion.
It is essentially promoting the transformation of traditional mobile GPUs into "AI graphics processors". Leveraging Neural Fusion, games can reduce power consumption, extend battery life and maintain stable high frame rate performance while maintaining peak performance. This native AI-driven technology can be used for neural processing, super-resolution and frame generation, aiming to revolutionize the way scenes are rendered and fused, and help the large-scale deployment of stronger visual effects on mobile terminals.
At present, Neural Fusion has become the first AI super-resolution technology to be implemented on mainstream game engines such as Unity and Unreal Engine.
Based on the brand new Adreno GPU, mobile phones equipped with Qualcomm Snapdragon 8 Gen 6 Super Ultra can now shoot 8K 60fps videos.
NPU: Supporting On-Device Deployment of 30B Large Models
The important new component of the new generation Hexagon NPU is Element Accelerator.
Qualcomm clarifies that it is designed for Transformer workloads used by generative AI and agents.
It will cooperate with the existing vector computing unit and scalar computing unit to specially accelerate key operations in large models. In other words, this generation of Hexagon is further tilted towards LLM/Agent at the architectural level.
At the same time, the new generation Hexagon NPU also expands the maximum capacity of the Large Shared Memory by 50%. This can keep context and KV-cache data inside the NPU, which can also reduce the number of DDR accesses, lower memory read latency, and free up valuable memory bandwidth for other workloads.
This means that even if the on-device model needs to process longer context windows, more complex reasoning, more tool calls and concurrent tasks, the agent experience can still respond in a timely manner.
In addition, for INT4 models, its prefill speed is increased by 50% to improve prompt processing and context understanding speed.
In addition to INT4, Qualcomm also reserves a relatively complete range of precision options for developers, supporting INT2, INT8, FP8 and FP16 precision, so that developers can strike a balance between performance, memory, model quality and power consumption.
One More Thing
Cristiano Amon, CEO of Qualcomm, clearly stated at the launch event that a new moment of transformation for mobile technology has arrived, and this time it is agent-centric.
And this is not something in the future, but something happening right now. At least, we are already at the beginning of this transformation.
Smartphone vendors including Xiaomi, vivo, REDMAGIC, REDMI, OPPO, OnePlus, Motorola, iQOO and Honor will take the lead in adopting the Snapdragon 8 Gen 6 Super Ultra and Snapdragon 8 Gen 6 Ultra mobile platforms in their flagship products.
Xiaomi also brought its upcoming new 18 Pro device to the launch event.
In the next two days, Qualcomm will release a series of new products around Personal AI, covering wearables, PCs and other devices.
QbitAI will also keep following up, and interested readers can stay tuned~
This article is from the WeChat Official Account "QbitAI", author: Yuyang, published with authorization from 36Kr.