From running AI models on smartphones to enabling real-world deployment of AI agents, Qualcomm's new story in the AI era | Focus Analysis
Over the past year, AI programming and office tools have rapidly gained popularity on PCs, but the AI experience on mobile devices is still dominated by single-turn Q&A and image processing. What Qualcomm aims to prove this year is that on-device AI can also be integrated into more life and work scenarios.
At the 2026 Snapdragon Summit, Qualcomm, together with Stepfun, Wulianhuo Technology and Jiangbo, demonstrated a typical case: a complete workflow powered by the on-device 30B-MoE model. Using only the local model, the AI can read an email, extract itinerary information, sync the calendar, recommend flights and hotels, and then draft a reply.
The highlight of this demo is that it is a task requiring understanding, planning and continuous execution, which is completed entirely on the end side.
Cristiano Amon, President and Chief Executive Officer of Qualcomm Technologies, Inc.
If Qualcomm was mainly proving two years ago that large language models can be deployed on mobile phones, this year it aims to further demonstrate that the models deployed on devices can handle more complete work tasks.
In 2024, Qualcomm's self-developed Oryon CPU was first equipped on mobile phones, showing the feasibility of on-device large language models; in 2025, the 5th-generation Snapdragon 8 Elite introduced Personal Scribe and personal knowledge graph, embedding the agent capability directly into the product roadmap.
Up to this year, the new trend in the AI industry has shifted to the code and office application competition led by Claude Code. Now Qualcomm is promoting its flagship chips to support 2nm dual-flagship configuration, 30B-MoE deployment and more complete task workflows, which means it can jointly expand new application scenarios with model and infrastructure vendors.
"Agent AI is not just an added function for end devices. It actually creates a brand new workflow for end devices," Cristiano Amon, President and Chief Executive Officer of Qualcomm Technologies, Inc., said in his keynote speech.
But when AI evolves from a single function to a full workflow, Qualcomm needs to do more than just make the next-generation chips faster. Agents require stronger local capabilities, and also rely on rapidly updated models, application interfaces and cross-device collaboration.
Entering 2026, a common trend among chip manufacturers is that their capability boundaries are expanding, and so are their competition boundaries. For Qualcomm, this expansion is not only from mobile phones to PCs, AR glasses, automobiles and data centers, but also includes complementing capabilities around model adaptation, software tools and cross-device collaboration.
From Deploying Models on Phones to Enabling Sustained Agent Operation
If there is only one keyword to summarize this year's Snapdragon Summit, "Agent" will definitely rank first. All updates from Qualcomm are reallocating computing resources in mobile phones for more complex AI tasks.
For example, both the 6th-generation Snapdragon 8 Elite and the 6th-generation Snapdragon 8 Super Elite dual flagship chips adopt 2nm process. For devices equipped with the Super Elite version, Qualcomm not only boosts the CPU frequency to 5GHz, but also assigns exclusive computing channels for different types of AI work: the CPU is responsible for tool calling and task scheduling, the NPU undertakes model inference, and the low-power perception unit processes background information.
The underlying change behind this is that AI no longer only works when the user asks a question, but also retains information in the background, waits for the right moment and continues unfinished tasks. In the past, a single Q&A or image editing task was a one-time call; agents need to read the context repeatedly, call tools and switch between different tasks.
Therefore, cache, memory bandwidth and low-power perception have become particularly important — at the very least, mobile phones should not keep overheating and draining power quickly just because a resident agent is running.
This is not a direction that emerged only this year. The 5th-generation Snapdragon 8 Elite released in 2025 has already emphasized Personal Scribe, personal knowledge graph and low-power perception hub.
This year, Qualcomm's upgrades in cache, NPU shared memory and perception capabilities are more like continuous reinforcement of the existing direction, enabling more information to be retained, and making task switching and continuous operation more reliable.
However, this does not mean that all AI computing is kept locally, but that end devices take over the more valuable part of the workload. Mobile phones hold personal information such as calendars and emails. Qualcomm's new dual Micro NPU architecture delivers 85% performance improvement while reducing power consumption by 20%, which directly enhances the Personal Scribe capability, bringing the advantages of local low-latency processing, offline availability and less data upload; complex tasks can still call cloud resources when needed.
Now, larger-scale models are starting to integrate with complete local office workflows. Compared with 2024, when models of up to 7B parameters could be deployed on end devices, Qualcomm has reached cooperation with Stepfun at this summit. Now, the 30B MoE model can complete email understanding, itinerary planning and calendar synchronization on the reference design, which is a tangible leap forward.
Chris Patrick, Senior Vice President and General Manager of Mobile Business, Qualcomm Technologies, Inc.
More complex AI applications can now be completed by the division of labor of multiple models, rather than relying on a single larger model. This creates more space for entrepreneurs in the application ecosystem.
For example, in the AI podcast application demonstrated by AI company Deepgram, the speech model on NPU is responsible for transcribing and distinguishing different speakers, and the Gemma model on GPU understands topics and generates follow-up question suggestions — it can clearly identify who is saying what while judging what to ask next. This makes "heterogeneous computing" a scenario that users can easily understand. However, the staff also mentioned that the application needs to connect to the internet for retrieval and verification, so the whole application is not completely offline.
AI smart glasses are also a typical product category. The changes over the past year can also illustrate how chip manufacturers like Qualcomm help new AI hardware gradually break through the upper limit of capabilities. At AWE in June 2025, the upgraded version of Qualcomm AR1 chip, Snapdragon AR1+ Gen 1, has demonstrated local operation of Llama 3.2-1B on prototype glasses. Qualcomm emphasized that the chip is smaller and consumes less power to adapt to the constraints of temple arms and batteries; this year, Qualcomm further provides multi-modal 1-bit model support for AR1 and AR1+.
Qualcomm is also gradually expanding its support for the developer ecosystem. The Qualcomm START program launched this year provides AR1+ modules, software platforms and reference designs, which lowers the integration threshold for non-technology manufacturers. It shows that R&D support is increasingly tailored to the needs of specific product categories, rather than just shrinking mobile phone chips for other devices.
In addition, Qualcomm has also launched the brand-new Adreno Neural Fusion, which enhances mobile gaming and realizes intelligent graphics rendering with AI, bringing a more immersive gaming experience, richer visual effects and longer gaming battery life.
Qualcomm's new generation Adreno
But what the chip can support is not the same as what the agent can ultimately complete for users. Meng Pu, Chairman of Qualcomm China, told 36Kr that personal agents that provide continuous cross-device services may not be steadily and gradually implemented until 2027-2028.
Chip Manufacturers in the AI Era: Faster Iteration, More Diversified Challenges
Qualcomm is now fighting on multiple fronts in the AI era, and the pressure is not small. Part of the pressure comes from the attraction of new markets, and the other part comes from the need to hold on to its existing market.
In the mobile era, the foundation of Qualcomm's success lies in its two business lines: communication patents and chip platforms. In the large model era, the huge demand led by data centers has brought massive revenue to NVIDIA — in the second quarter of fiscal year 2027, NVIDIA's data center revenue reached 89 billion US dollars.
Qualcomm has not missed this market either. Previously, Qualcomm set a target of 40 billion US dollars in non-mobile revenue by fiscal year 2029, and is advancing its data center business. The motivation is easy to understand: on the basis of securing its position in the mobile phone market, it has to find new growth engines.
But before pursuing new growth, Qualcomm may first need to adapt to the model market that iterates faster than the mobile phone replacement cycle. "In the past, after a smartphone completed adaptation, it might not need any adjustment for 9 or even 12 months; but now, large models iterate every 3 months or even shorter," Meng Pu, Chairman of Qualcomm China, told media including 36Kr.
The rapid iteration of model versions and capabilities does not mean that the underlying computing structure will change completely at the same speed. The more accurate challenge is: on a relatively stable computing foundation, use software scheduling, configurable capabilities and continuous verification to support rapid model iteration.
A deeper change is that cooperation has moved forward from the adaptation after model release to the model design stage. The division of labor of the four-party cooperation demonstrated in 2026 illustrates this point: Qualcomm provides the computing platform, Stepfun provides the model, Wulianhuo optimizes inference and scheduling, and Jiangbo Long handles storage collaboration; the AI Hub tries to reuse optimization results and lower the deployment threshold.
The new competition metric is: who can deliver usable experiences faster and at more controllable costs after new models are released.
On-device agents are opportunities that the whole market is chasing. In September, almost all new products of mainstream manufacturers cover on-device AI and agents. For example, on September 15, MediaTek released the Dimensity 9600 series, and the 9600 Pro also adopts 2nm process, emphasizes low-power resident perception, and claims to support 30B-MoE on-device operation.
More complex competition also comes from the customer side: while purchasing Qualcomm chips, customers are also trying to master key capabilities on their own. For example, Xiaomi's Ring O1 has entered mass production, and Apple has launched its self-developed C1 modem on iPhone 16e. When entering the data center market, Qualcomm also cannot avoid the situation where customers and competitors overlap. In addition to NVIDIA, Google's self-developed TPU is also competing for training and inference workloads — and Google is also an important partner of Qualcomm on the end device side.
When competition occurs in multiple markets at the same time, Qualcomm must also answer a practical question: the premise is that customers can afford these new capabilities, but how to apply them efficiently in real scenarios?
Meng Pu mentioned: "With the advancement of technology, more and more technologies need to be integrated into chips, and the cost will therefore get higher and higher." In his view, one product can show the upper limit that current technology can reach, while the other focuses more on efficiency to cope with high costs and support terminal manufacturers. Even if technical capabilities continue to advance, cost control is still a key barrier for customer adoption.
At least, there are still multiple challenges to make agents run stably on mobile devices. "The agent can't perform wrong tasks, can't occupy too much memory, can't increase the power consumption burden, and at least can't run slower than manual operation," an insider in the semiconductor industry summarized.
Now the chip industry is entering a multi-front melee. Qualcomm needs to defend its existing territory and explore new markets, and must also become a long-term collaborator that connects large model vendors and AI applications. For Qualcomm, the success of the next round does not only depend on how large a model it can deploy on end devices, but also on whether it can turn the constantly evolving models into products that users are willing to use repeatedly.