HomeArticle

Integrate 100-billion-parameter large language models into laptops, Windows is finally set to strike back against Mac.

爱范儿2026-09-08 07:53
All want to make PC the runtime environment for AI.

The concept of AI PC has been hyped for two years, and finally, someone has moved the "local large model operation" from the demonstration stage into people's backpacks.

Lenovo launched a full range of AI terminals at Lenovo Innovation World 2026 during IFA 2026 in Berlin, including smartwatches, mobile phones, tablets, laptops, monitors and other products.

The most important product among them is undoubtedly the Lenovo Yoga Pro 9n (domestic model: YOGA Pro 15 Spark) jointly launched by Lenovo and NVIDIA, which is a mass-produced Windows laptop capable of running large models with hundreds of billions of parameters locally.

This laptop weighs 1.65 kg, is 16.7 mm at its thinnest, and is equipped with up to 128GB of unified memory. According to NVIDIA's specifications for RTX Spark, this configuration can already run large language models with 120 billion parameters and a maximum context of 1 million tokens locally.

In the past when we discussed AI PCs, they were either professional hosts placed on desks, such as DGX Spark and Mac Studio, or Mac Mini that made people scramble to get them at all costs.

For a laptop to be a qualified AI PC, it seems that it either needs to push the hardware to the extreme as much as possible, with large memory and top-level graphics cards, while keeping the size and heat dissipation of the computer under control.

Or in terms of software, it used to be acceptable to install OpenClaw without obstacles, and to support the use of various AI Agents to the maximum extent. Just like AI phones, the more apps that AI can operate, the more "AI-capable" the device is, and the same goes for computers.

But now the AI PC track has gone through a complete cycle: the concept becomes popular, manufacturers compete on parameters, and users begin to question the actual value of these products. We also start to wonder what exactly these computing powers can do on our computers.

It may be more than enough to handle function-level tasks such as real-time translation and conference noise reduction, but complete large model inference and slightly more complex AI workflows still need to be processed back to the cloud.

Local large models are more often placed on the demonstration stage, rather than on our desktops.

The Yoga Pro 9n launched by Lenovo this time takes a different path. It is deeply integrated with NVIDIA's end-side AI architecture, equipped with the NVIDIA RTX Spark superchip, and integrates the 20-core NVIDIA Grace CPU and the Blackwell RTX GPU with 6144 CUDA cores through NVLink-C2C interconnection, delivering 1 PFLOPS (FP4 precision) of AI performance.

A Windows laptop now also has the potential to become an Agent host.

Unified Memory is the new trend of AI PC

In traditional PCs, CPU memory and GPU video memory are independent of each other, and large model parameters are repeatedly transferred between the two. The capacity of video memory often directly determines the maximum size of the model that can be run.

In the past, discrete graphics cards were almost standard for high-performance computers, but no matter how strong the peak computing power of discrete graphics cards is, the video memory of consumer products is usually only a dozen or 20 GB, which is not enough to hold a slightly larger model.

The unified memory architecture of NVIDIA RTX Spark allows the CPU and GPU to share the same physical memory pool, with up to 128GB available for both at the same time. Model weights, input data and intermediate results only need to be saved once; large models of hundreds of billions of parameters can also run locally on a laptop.

The reference value given by NVIDIA is a large language model with 120 billion parameters and 1 million token context.

Developers can continue to use the familiar CUDA technology stack, no data needs to be uploaded to the cloud, and project documents, codes and materials are all stored locally. The platform can also coordinate multiple AI agents in the background to automatically complete complex creation tasks across applications.

It is worth mentioning that Lenovo's X Power cooling architecture squeezes 80W TDP into a body as thin as 16.7mm, equipped with a 15.3-inch 2.5K 165Hz PureSight Pro OLED screen; based on the RTX platform, the Yoga family also obtains advanced gaming capabilities.

Developing AI applications, editing high-resolution videos, and rendering 3D scenes can all be completed on this single machine.

Luca Rossi, Executive Vice President of Lenovo Group and President of the Intelligent Devices Group (IDG)

The Lenovo Yoga 9n 2-in-1 on the same platform packs this computing power into a 360-degree flip creative laptop: up to 64GB of unified memory, supporting dual-surface handwriting on "screen + trackpad".

2 in 1 refers to Lenovo's dual-mode 2-in-1 laptop

Lenovo is not the only player betting on unified memory. Not long ago, Apple's newly updated Mac mini and Mac Studio have also almost raised the upper limit of local AI computing performance to a new level.

The Mac mini comes with two chips, M6 and M5 Pro, with up to 64GB of unified memory, and supports connecting multiple machines into a cluster through Thunderbolt.

The Mac Studio offers M5 Max and M5 Ultra. The M5 Ultra version supports up to 512GB of unified memory with a bandwidth of 1.2TB per second, which according to Apple, is enough to carry models with tens of billions or even hundreds of billions of parameters.

At a special event of WWDC26 in June this year, Apple also demonstrated the operation of this kind of cluster on site: 4 Mac Studios are interconnected through Thunderbolt 5, and run the open weight model Kimi K2.6 with trillion-level parameters via LM Studio based on MLX Distributed. Official data shows that the inference performance of the 4-machine cluster can reach up to 3 times that of a single machine.

For Apple's laptops, the 128GB unified memory MacBook Pro (M5 Max) can also run the 120B parameter open source model gpt-oss-120b locally. Actual tests show that the speed can reach about 79 token/s at most, but the premise is 4-bit quantization, and long context will significantly occupy memory.

Putting the update directions of the two PC manufacturers together, it is obvious that the bottleneck of end-side AI has shifted from pure computing power to the competition of "whether the model can be loaded and whether the data can be transferred smoothly". Apple uses unified memory plus Thunderbolt cluster to cover the entire product line from laptops to desktops.

The significance of Lenovo and NVIDIA's cooperation this time is that they have brought the same capability into the Windows + CUDA ecosystem for the first time, and pushed the upper limit of local context to the order of 1 million tokens with FP4 precision and unified memory.

Windows is also competing for the Agent host market

But in this competition, there is another easily overlooked change: Windows has entered the game. When we look at this matter from a broader ecological competition perspective, the significance of Lenovo Yoga Pro 9n is far more than just "another powerful AI PC".

For a period of time in the past, devices that can stably carry local large models, long contexts and multi-agent workflows were mostly discussed around Mac and Linux operating systems.

With unified memory and MLX, a small "desktop device" can make Mac Mini sell out; Linux is the base camp for NVIDIA, workstations and developer toolchains.

Personal AI supercomputers like DGX Spark are built on the Linux software stack, and a large number of CUDA development and model deployment tools are first matured around Linux.

While Windows has always been the most widespread personal computer platform, it is rarely regarded as the home of local large models and Agents.

Now, as Lenovo integrates the RTX Spark capabilities into mass-produced laptops and brings scalable small supercomputing nodes into commercial scenarios, Windows also has the realistic foundation to carry hundreds of billions of parameter models and local Agents.

Introduction page of NVIDIA RTX Spark chip, many Windows computers are equipped with this superchip

For people who have developed AI applications around CUDA, RTX Spark is backed by years of accumulated development tools, software libraries and developer groups of CUDA; existing knowledge and development experience can continue to be used, which also gives Windows an extra software advantage in the competition for personal Agent hosts.

While Mac is becoming a local AI workstation for developers and creators, Windows obviously does not want to miss this market.

Beyond laptops, everything is serving for Agents

After expanding and unifying computing power and memory, the next question is: who are these capabilities prepared for? Lenovo's answer is AI Agent.

When AI evolves from a "responsive assistant" to a "collaborator that can plan independently and work continuously", its requirements for devices have also changed: tasks run for hours, contexts need to be retained for a long time, and data and permissions must be stored locally.

Lenovo's new commercial product ThinkCentre X Ultra (domestic model: ThinkCentre X Tiny) is designed for this demand. In a 1.6-liter body, it is equipped with AMD Ryzen AI Max+ PRO 495 series processor and up to 128GB of unified memory.

Different from DGX Spark which only supports Linux system, Lenovo provides pre-configured AI tools and workflows for Windows and Linux for this supercomputing box

What's more interesting is its expansion method: up to 4 devices can be interconnected into a unified computing platform to run larger models, longer contexts and multi-agent workflows.

A department can start with one device, and add more devices gradually when the business grows and the model becomes larger. From "one computer per person" to "one for AI and one for me", the desk is becoming the smallest AI computing node.

In addition to "bringing hundreds of billions of parameters into laptops" and this personal supercomputer benchmarked against NVIDIA DGX Spark, Lenovo's product lineup this time is very rich, and we have simply sorted out the following information:

Two concept devices: Project Aeroblade uses AirJet solid-state active cooling developed in cooperation with Frore Systems to replace fans, achieving full fanless performance in a body less than 10mm thick and weighing less than 1000 grams; Project Swan brings the rollable screen to commercial laptops, which is 14 inches when closed and 17 inches when unfolded.