DeepSeek, Qwen and Zhipu AI have appeared in succession, and PC vendors have finally obtained the ammunition they have been waiting for.
No new graphics cards, no benchmark bombardments, AMD's opening keynote at IFA 2026 focused on one single thing: Personal AI.
Along with three new products that underpin it: the Ryzen AI Max 400 (codename Gorgon Halo) platform, the first reveal of the next-generation HP ZBook laptop, and a 96-core, full-liquid-cooled "desktop supercomputer" called the Threadripper Halo Station.
Jack Huynh, Senior Vice President of AMD and General Manager of Computing and Graphics, spent the 45-minute speech tracing from the Artemis II lunar flyby, surgical robots to Finland's LUMI supercomputer and PlayStation 5 Pro, reviewing AMD's full computing power landscape.
He then turned the topic to everyone's personal computers: when AI plans, invokes tools and modifies results all day long like an employee, where exactly should all these computations be run?
Jack Huynh, Senior Vice President of AMD and General Manager of Computing and Graphics
AMD's answer is Personal AI. The PC handles tasks it can complete locally first, keeping tasks involving sensitive data and personal context on the device as much as possible; when larger-scale capabilities are needed, it then calls on the cloud and data center resources.
Following this vision, AMD has covered everything from PCs supporting 192GB of unified memory to desktop workstations, and partnered with Microsoft and SUSE to supplement Windows development and enterprise deployment workflows. Different products and partners all seem to be answering the same question: how to make AI start working on the devices closest to the user first.
DeepSeek, Qwen, GLM, gpt-oss and Laguna also appeared on stage, serving as new performance measurement benchmarks in this keynote.
The Qwen3.8-Flash-Next, released just over a week ago, can run natively on AMD's new platform, and the GLM-5.3-Flash with a total of 320 billion parameters can also be loaded into the unified memory. AMD used these open-weight models to prove that from regular AI PCs to desktop workstations, every piece of hardware can indeed run corresponding AI tasks.
For these PC chip makers, they used to compete on game frame rates, but now the competition has shifted to comparing how large a model their devices can run locally.
"AI is the new electricity", but the cost is getting unaffordable
Jack condensed the evolution of generative AI over the past few years into four stages: it first learned to converse, then gained visual and speech capabilities, and later developed reasoning abilities; by 2025, AI has evolved from a question-answering model to an Agent that can plan, call tools, check results and keep working autonomously.
This kind of transformation not only brings better user experience, but also leads to a sharp increase in workloads.
A regular ChatGPT conversation only consumes hundreds to thousands of tokens, while an Agent task like "help me plan a family trip to Italy" that needs to check calendars, book flights and adjust hotel reservations will consume hundreds of thousands or even millions of tokens in a single run.
According to figures given by AMD, consumer-level users will consume millions of tokens per day, creators will use tens of millions, and developers and entrepreneurs will start at 50 million tokens per day.
Macroeconomic data shows that over the past year, the global monthly AI token processing volume has surged from about 0.7 quadrillion (700 trillion) to 1.7 quadrillion (1700 trillion); by 2030, it is expected to reach 120 quadrillion (120,000 billion) per month.
At the same time, 93% of enterprises are overspending their AI budgets. Jack mentioned that for an active user generating 15 million output tokens per day, the cloud computing cost is about 300 euros per day; to pay for one person's AI usage for a whole year, it will cost nearly 100,000 euros.
"We can't just spend more money, we have to compute smarter." This line is probably the core theme of AMD's entire launch event, how to bring AI back from the cloud to local devices.
Models are getting smaller, PCs are getting more powerful
The premise that the local computing narrative holds is that local models are good enough compared to cutting-edge closed-source models, namely the model size gets smaller while its performance gets stronger.
The open-weight model GPT-OSS-120B released by OpenAI last August scored 80 points on the GPQA (graduate-level reasoning benchmark) and has become the first choice for local deployment for many developers.
A few months later, Alibaba's Qwen3.5-9B, with only 9 billion parameters, got a higher score on the same benchmark, with 13 times fewer parameters but better performance.
In recent months, the competition between open-source models and closed-source models has become increasingly fierce, from Kimi K3 competing with Fable 5, to DeepSeek, GLM, and even Flash models starting to match the performance of the Claude Opus series of models.
AMD's roadmap is also developed along this curve: the current Gorgon Point processor, namely the Ryzen AI 400 series, can run 24-billion-parameter models locally; the Strix Halo with 128GB unified memory can run 200-billion-parameter models.
AMD also presented on-site comparisons of the performance of these local models: the open-source model Laguna S 2.1, running entirely on Strix Halo, outperformed the cloud-based Claude Sonnet 5 (70.3 points) in a software engineering benchmark.
For 10 million output tokens, the cloud cost of Claude Sonnet 5 is about 90 euros, while the local cost is zero
Moving computing back to personal computers instead of placing it in data centers is AMD's vision for local computing.
After local computing solves the problems of speed and cost, privacy and control also become a sensitive aspect for Agents. AMD hopes that these data will stay with the user first, and the user can decide what content can leave the device and where it will be sent.
The last part is personal context, which Jack regards as the fourth layer beyond applications, models and computing power. The same sentence "help me prepare for tomorrow" may mean client meetings and three pending matters for office workers, while for students it may mean exams and study group arrangements.
It is not enough for the model to know general world knowledge, it also needs to know what the person asking the question is going to face the next day. Local computing, privacy and control, personal context - these are the three key words of Personal AI that AMD is betting on.
In terms of specific products, AMD introduced Gorgon Halo, whose official market name is Ryzen AI Max 400 series. Its unified memory is increased from 128GB to 192GB, and the maximum parameter of models that can run locally is raised from 200 billion to 300 billion.
At the event, AMD also demonstrated GLM-5.3-Flash from Zhipu AI, a 320-billion-parameter model that can run locally on Gorgon Halo, and showed that on the same benchmark, the performance of the local model is even better than the cloud-based Claude Fable 5; as a comparison, running Fable 5 to generate 10 million output tokens costs about 500 euros.
Lenovo also launched the ThinkCentre X Ultra, a flagship device at IFA 2026, equipped with Gorgon Halo. It is small enough to be placed on a desk, and up to four units can be combined to form a cluster.
In addition, Jack also unveiled for the first time the next-generation HP ZBook codenamed Sundance: it has 192GB of unified memory, and on site, it used a single prompt to generate an entire 3D world in real time locally, and then ran *Microsoft Flight Simulator* smoothly on it.
"Create worlds during the day, play games at night. One hand holds the real world, the other holds the world I imagine."
Gorgon Halo is still a compact personal computer. Near the end of the keynote, Jack expanded the concept of "Personal AI" to another form factor.
The newly debuted Threadripper Halo Station integrates a 96-core Threadripper PRO and a data-center-class AMD Instinct MI350P into a single desktop workstation.
The demo unit on site is equipped with two MI350P GPUs, and reserves space for expanding to four; the maximum system memory reaches 2TB, the total HBM3E capacity under the 4-GPU configuration is up to 576GB, and the whole machine adopts full liquid cooling.
AMD claims that it can run models with over 1 trillion parameters, and calls it "the world's most powerful workstation".
This device extends the Personal AI concept AMD advocates from a single PC to a shared node in the office. A single developer can use it directly, and the whole team can also call its resources together; tasks that cannot be handled by personal PCs do not necessarily have to be sent to the public cloud that charges by the token immediately.
The ecological puzzle of local AI
Having hardware that can load the model does not mean that the Agent can already run safely on the computer.
It also needs to perform operations across applications, services and files, and the operating system must be easy to use. If the computing power is strong but the system only runs Ubuntu of the Linux distribution, it will be very difficult for ordinary users to get started. In addition, the software must also control what the Agent can view, what it can modify, and when user authorization is required.
Pavan Davuluri, Executive Vice President of Microsoft Windows and Devices Business, introduced on stage that Windows is being adapted for unified memory access and workload scheduling, so that large models can make full use of the memory and GPU of Gorgon Halo.
Microsoft's previously announced Microsoft Execution Containers aims to provide an isolated execution environment for Agents to protect users, their data and devices; this capability is still in the early preview stage at present.
Microsoft also announced Project Zenith, an out-of-the-box Windows experience for developer-grade hardware. It requires devices to have at least 64GB of unified memory, pre-installed with VS Code, WSL, GitHub Copilot CLI, PowerShell and a set of developer-focused settings, and will be first available on Ryzen AI Halo devices.
He also demonstrated on site the performance of Alibaba's new 125-billion-parameter model Qwen3.8-Flash-Next running on Gorgon Halo. From small tasks such as "help me write a function", to local forensic analysis of sensitive files, to generating a designer portfolio website with a single sentence, the whole process is completed locally.
Pavan gave this vision a name: PC will become the home of "unmetered intelligence".