HomeArticle

Does AMD Want to Bring "Token Freedom" to Everyone?

36氪的朋友们2026-09-07 08:54
Bet on personal computing.

"AI PC is the infrastructure for personal computing, not a substitute for cloud-based AI."

In January, Lisa Su stated at CES that AMD would bet big on AI personal computing, and the solution AMD presented at that time was the 128GB Ryzen AI Halo.

Eight months later, AMD decided to further boost the configuration of personal computing infrastructure, officially launching the new-generation Gorgon Halo chip, which increases the unified memory from 128GB to 192GB, allowing users to deploy 300-billion-parameter models locally.

AMD's upgrade strategy is directly targeted at NVIDIA, whose DGX Spark chip launched for personal AI computing has 128GB of unified memory and only supports local deployment of models with a parameter scale of up to 200 billion.

At IFA (Berlin International Consumer Electronics Show) in early September this year, Jack Huynh, Senior Vice President of AMD and General Manager of Computing and Graphics, further extended Lisa Su's views in his keynote speech, emphasizing that AI is entering the "Personal AI" era. Intelligence will no longer only reside in cloud data centers, but be deployed on the computers at your fingertips, the laptops in your backpack, and even the workstations on your desk that can run trillion-parameter models.

If data center GPUs are NVIDIA's core battlefield, then in the server CPU market, AMD has strongly taken half of the market share — EPYC captured 46.2% of the share in terms of revenue in the first quarter of this year.

Against this backdrop, the personal AI market focused on local computing will become a new battlefield for upstream chip companies to tap into new revenue streams. This is why Lisa Su has mentioned on multiple occasions when talking about AI that "we are only in the third inning", and she also predicted that the number of global active AI users will rise from 1 billion to 5 billion by 2030.

Tracking data from Counterpoint also provides guidance for the "Personal AI" chip war: the data shows that the global shipment penetration rate of high-end AI PCs will reach approximately 59% in 2026, a sharp jump from 39% in 2025.

Therefore, Counterpoint also defines this year as the "crossing year" when hardware catches up with software.

Evolution Curve of Token Processing Volume

During the speech, Jack Huynh divided "Personal AI" into three parts: local computing, privacy and control, and "personal context". He also quantified the forecast of user scale to the Token level.

"The monthly processing volume will rise from about 0.7 quadrillion to 1.7 quadrillion, and is expected to reach 120 quadrillion by 2030. Based on 15 million output Tokens per day, the pure cloud cost is about 300 euros per day, and nearly 100,000 euros per year."

In other words, moving part of the computing back to local devices will turn "smart computing" into "more affordable computing".

To this end, AMD has formed a "Personal AI" alliance, extending from the Gorgon Halo chip to products such as the Lenovo ThinkCentre X, HP ZBook, and the more computationally aggressive Threadripper Halo Station.

HP ZBook Codenamed Sunday

The narrative of "Personal AI" can be established by shifting work that can only be advanced via cloud subscriptions to local devices. The improvement of hardware specifications is only one aspect. The other aspect depends on the evolution of models.

"AI is getting smaller, faster, and more powerful at the same time," Jack Huynh said, and this efficiency improvement will be directly applied to PCs. However, he also emphasized that local deployment does not mean an inferior choice.

In software engineering benchmarks, the open-source model Laguna S 2.1 deployed locally on Strix Halo outperforms the cloud model Claude Sonnet 5; GLM 5.3 Flash running locally on Gorgon Halo also beats the leading cloud model Claude Fable 5.

In terms of economy, 10 million output Tokens of Sonnet 5 cost about 90 euros. Running the same scale in high-usage scenarios on Gorgon Halo costs about 500 euros if billed by Token volume, but local deployment does not involve pay-per-use billing. In other words, after paying the hardware cost, users will have continuous "Token freedom".

In addition to the booths of OEM partners such as Lenovo, Microsoft also provides support at the software level.

Pavan Davuluri, Vice President of Microsoft and Head of the Windows and Devices Division, mentioned in a discussion with Jack Huynh

The Agent evolves from "understanding intent" to "taking actions across applications, files and services", which is based on trust, security and user control.

At the same time, Pavan also announced an out-of-the-box Windows programming experience called Project Zenith, which is pre-installed with VS Code, WSL, GitHub Copilot CLI and PowerShell, and explicitly "supports AMD Ryzen AI Halo out of the box", allowing developers to run large models and complex AI workflows locally.

However, the Threadripper Halo Station workstation is the ultimate implementation of "bringing computing power right next to individuals".

The Threadripper Halo Station is equipped with a 96-core Threadripper PRO processor, 2 Instinct MI350P accelerators (scalable to 4), up to 2TB of system memory, up to 576GB of HBM3E, all adopting a liquid cooling solution, supporting local deployment of trillion-parameter models.

On-site Photo of Threadripper Halo Station

From AI PCs based on Gorgon Halo to Threadripper Halo Station workstations, the Token freedom that AMD strives for for "Personal AI" actually provides users with a step-by-step choice.

At the same time, as models, tools and infrastructure will all evolve, developers also have "freedom of choice" — choose local deployment if the workload is suitable for local operation, and use the data center if it is more suitable for data center operation.

"Every great computing era has empowered people with capabilities that they did not have in the past. And now, AI is empowering us with a brand new capability," Jack Huynh said.

The following is the edited full transcript of Jack Huynh's speech (abridged without altering the original meaning)

Good morning, IFA.

For decades, we have been continuously making computers faster, smaller and more powerful. We have placed computers on our desks, and also put them in our pockets. Now, we are integrating intelligence into every place. We are endowing computers with the ability to understand — understand what we are doing, what we want to achieve, what is most important to us, and finally work collaboratively with us. And this will change everything.

At AMD, we build the computing capabilities that support all these possibilities, extending from the smallest devices all the way to the world's largest supercomputers. We have been solving the most difficult problems in the computing field, and we always adhere to a simple concept: technology is most valuable when it truly puts more capabilities in people's hands. Today, I want to show you what will happen when this capability becomes more personalized.

From the edge of space, to the operating room, to the future of our planet, AMD technologies are powering all of these. Computing not only changes what machines can do, it also changes what people can do.

AI is the electricity of the new era. Just as electricity was eventually integrated into daily life, intelligence will also be integrated into the tools we use every day. AI is entering every aspect of our lives, helping us discover new drugs, see things that were invisible in the past, and turn ideas into reality.

Every generation leaves a more powerful set of tools to the next. PCs put computing power in the hands of millions of people; the internet makes global knowledge accessible; smartphones put computing in our pockets. And now, AI is endowing computing with intelligence. The question is no longer just "what can computers do", but "what can we do with computers".

AI is developing at an extremely rare speed in human history. A few years ago AI learned to converse, a year later it learned to "see", then it learned to reason, and in 2025 it entered the Agentic stage — it no longer just answers questions, but becomes a system that can take actions: it can plan, use tools, check its own work and execute continuously. One Agent helps one person do more work, and a group of Agents work in parallel, so the multiplier effect keeps expanding.

But this new leverage comes at a cost. Every time productivity doubles, computing demand doubles. Agentic AI is driving the growth of Token consumption at an unprecedented speed. A standard ChatGPT interaction only uses hundreds to thousands of Tokens; while an autonomous Agent planning a trip may use hundreds of thousands or millions of Tokens, and this happens all day long.

The standard models we use every day process about 0.7 quadrillion Tokens per month, rising to 1.7 quadrillion, and this is before the emergence of Agents. By 2030, the monthly volume is expected to reach 120 quadrillion. The gap is huge. We need more computing, and we also need to rethink where this computing runs.

93% of enterprises have exceeded their AI budgets, and Agents will only further increase demand. We can't just spend more money, we have to compute more intelligently. Based on 15 million output Tokens per day, the cloud cost is about 300 euros per day, nearly 100,000 euros per year. At the same time, chips are getting smaller and more powerful, and personal devices have reached the performance level that only data centers had in the past. The next generation of computing will be defined by intelligence that is closer to you — more personalized, more efficient, and more capable.

We are entering the era of Personal AI. Personal AI understands your context, is tightly connected to your data, and expands your imagination and achievements. It includes three core parts: local computing, privacy and control, and personal context.

Let's start with local computing. Last August, OpenAI released GPT-OSS, a 120-billion-parameter open weight model, which scored 80 points on GPQA; a few months later, Qwen 3.5 with 9 billion parameters surpassed it on the same benchmark — 13 times fewer parameters, but better results. AI is getting smaller, faster and more powerful. Strix Halo uses 128GB of unified memory to run up to 200 billion parameters locally; Gorgon Halo increases the unified memory to 192GB, supporting up to 300 billion parameters running locally. What used to require cloud subscriptions can now be done directly on your computer.

Local deployment does not mean an inferior choice. In software engineering benchmarks, the open-source model Laguna S 2.1 running on Strix Halo outperforms the cloud model Claude Sonnet 5; GLM 5.3 Flash running locally on Gorgon Halo also beats the cloud model Claude Fable 5. In terms of economy, 10 million output Tokens generated by Sonnet 5 cost about 90 euros, while the same scale in high-usage scenarios on Gorgon Halo costs about 500 euros — but there is no pay-per-Token billing for local use, you can use it once or a thousand times. On Strix Halo, 10 AI Agents can work in parallel locally, turning the laptop into a team that collaborates with you.

Right behind me is the Gorgon Halo Box, which is built for developers to explore new categories of local AI products, and will be available in Europe soon. Our OEM partners are also very excited.

Yesterday, Lenovo launched the Gorgon Halo-powered ThinkCentre X in Berlin. It is small enough to be placed on the desk, and can run 300-billion-parameter models locally. We also built a laptop for the first time with Gorgon Halo, codenamed Sunday, which has been under development for several years — the new generation HP ZBook, with 190GB of unified memory. The entire 3D world can be created locally with one instruction, and flight simulation can also run smoothly. This is Personal AI in your hands.

Both models and devices are ready, and cutting-edge technology is no longer out of reach. But Personal AI must also win trust, which leads us to privacy and control. Personal AI touches your most valuable assets — files, conversations, preferences. The closer intelligence is to you, the more important it is to protect this personalized intelligence. Sensitive work should stay on your device and be under your control; your context moves with the device you choose, and is shared only when you choose. Most AI today starts from the cloud, but imagine intelligence starts from your side: the PC handles what it can process locally, and connects to the cloud only when larger scale is needed, with the PC deciding where the workload runs.

The last component of Personal AI is personal context. The most powerful AI does not just know more about the world, it also knows more about you. The computing stack includes applications, models and computing, but Personal AI needs an extra layer of context — which tells who you are, what is important, and what you are striving to achieve. Without context, AI only gives a generic answer; with context, it gives the right answer tailored to you. Two people ask "help me prepare for tomorrow", one means a client meeting and three decisions, the other means an exam — the same question, but the needs are completely different.

We have been working with Microsoft for nearly 20 years, jointly pushing Strix Halo and Gorgon Halo forward continuously — stronger CPU, GPU, AI, performance per watt and security, AMD keeps delivering. Microsoft has launched Project Zenith, an out-of-the-box Windows programming experience that will support AMD Ryzen AI Halo out of the box, with at least 64GB of memory, allowing developers to run large models and complex AI workflows locally; we look forward to being the first to provide this experience on the Gorgon Halo Box.

Things that used to require the cloud can now run directly on the local PC. PCs are evolving from tools that respond to you to objects that understand you and help you complete tasks, which requires collaboration across chips, software, operating systems and applications.

We just discussed how AI is changing PCs, but what happens when a system can become any form needed for work? What if a PC does not have to be just a PC, if the system can provide developers with excellent computing power, and become a shared AI infrastructure when the team needs it, if modern data center-level computing power is brought right beside you? We have designed and built it.

Today we are sharing for the first time a brand new category — Threadripper Halo Station, the world's most powerful workstation, which can run AI models with parameters exceeding one trillion. It is equipped with a 96-core Threadripper PRO, 2 Instinct MI350P accelerators (scalable to 4), up to 2TB of system memory, up to 576GB of HBM3E, all with liquid cooling — it is almost your personal supercomputer.

We started with Personal AI, bringing excellent computing power to the desktop with Gorgon Halo, and then bringing supercomputer-level computing power to the workstation. Now we need to make this computing power open, accessible and ready for the world, which requires software and ecosystem support.

How to further expand the personalized capabilities of PCs and workstations? Workstations are not the end point, they are just the starting point. The real transformation is: developers' ideas start from one Ryzen AI Halo, enter a secure and supported production environment, and then expand to enterprise-level infrastructure as demand grows. Scaling up should not mean giving up choices — models, tools and infrastructure will all change, developers and enterprises should always hold the initiative. AMD has the same open DNA, we want to always stay open and never narrow the space of choices.

The bigger opportunity is to let hundreds or even thousands of AI systems — workstations, servers, data centers, edge devices — work collaboratively instead of operating independently. Expand from one system to multiple systems, keep open, and let all computing power collaborate seamlessly. An idea starts from one person