HomeArticle

The "Little Era" of On-Device AI

象先志2026-09-08 15:26
The great era will have to wait a while.

The start of this "mini era" was marked by a Mac mini buying spree. 

In late January 2026, an agent then named Clawdbot and later widely known as "Lobster" went viral, which subsequently triggered the Mac Mini buying frenzy.

People began to prepare a separate dedicated computer for AI that can execute tasks at any time.

Figure 1|The M4 Mac mini launched in 2024. Image source: Apple Newsroom.

Later, Apple has deliberately shifted its communication focus to agents. The new Mac mini no longer only highlights office and creative scenarios, but also emphasizes running intelligent agents 24/7; Mac Studio continues to target users who need to run large models locally by providing large memory configurations.

Xiaomi also recently showcased its AI Cube. Three Xuanjie chips are installed in an aluminum alloy cube, paired with 80GB of unified memory to deploy large models locally.

NVIDIA, which entered the market earlier, started delivering DGX Spark in October 2025. Lenovo, ASUS, HP and Dell have also launched complete devices based on GB10. Companies that originally sold laptops, workstations and small form-factor PCs are beginning to compete for a new type of order: users want to keep part of their AI computing power on their own premises.

Is edge AI about to usher in a big era? Hold on, it is still in the "mini era" for now.

Models are rewriting PC specifications

Just two years ago, the most typical configuration of an AI PC was only equipped with an NPU.

Microsoft set the threshold for Copilot+ PC at 40 TOPS NPU, 16GB memory and 256GB storage. Manufacturers introduced local AI with features such as real-time captioning and image processing. This set of standards can be applied to existing lightweight laptops, allowing users to use computers in their familiar way, with only a few additional functions supported by dedicated computing units.

But at the same time, starting with Apple's unified memory, users gradually began to deploy local models. In March 2025, Apple had put up to 512GB of unified memory into the M3 Ultra Mac Studio, and took running hundreds of billions of parameter models locally as its selling point. At that time, local deployment was only a choice for professional workers, because the Mac Studio, which cost tens of thousands of yuan, was far from the purchasing range of ordinary consumers.

This year, edge AI products continue to evolve. The Xiaomi AI Cube prototype is equipped with 80GB of unified memory, DGX Spark and a number of GB10 complete devices come with 128GB, and the AMD Ryzen AI Max 300 series supports up to 128GB; the new Mac Studio continues to offer up to 512GB, with memory bandwidth increased to 1.2TB/s.

It is no secret that memory is the top priority. Model weights, context cache and runtime will consume memory rapidly. Roughly calculated with 4-bit quantization, a 70B model requires about 35GB of memory only for weights, not including quantization metadata, cache and system overhead.

After deployment is completed, speed becomes the next threshold.

The 128GB unified memory of DGX Spark is paired with a bandwidth of 273GB/s. The former determines how large a model can be loaded, and the latter affects the data transmission speed during computing. When actually generating responses, model precision, context length and inference software will further affect the final results. Increasing memory alone cannot solve all subsequent problems.

Even if the model is already running, users may not be able to use it smoothly. Issues including how to download and update the model, how applications call it, what permissions are required before reading files, and who is responsible for maintenance when continuous operation fails, all need to be addressed. What users buy is a computer, but these links may not be completed by the same company.

Lenovo, ASUS and HP can adopt NVIDIA's GB10 platform to complete the design, sales and service of complete devices; Apple can adjust chips, systems and development tools at the same time. Xiaomi has installed its own chips in the prototype, trying to participate in defining how a personal AI device should perform computing.

These different paths reflect how much R&D investment each company is willing to make, and how much product decision-making power they want to grasp. Manufacturers that adopt off-the-shelf platforms can focus more on productization and delivery; those that design chips on their own need to take on R&D and verification investment first, in exchange for greater adjustment space for memory architecture, computing units and power consumption allocation. In the follow-up, sufficient product sales are required to dilute the cost of continuous iteration.

Self-developed chips are therefore a choice, not an entry requirement for all participants. Two machines that both use GB10 may have very similar underlying capabilities, but their installation experience, maintenance services and customer relationships can be controlled by different companies.

Figure 2|Industrial division of labor for personal edge AI. A single company can participate in multiple links, and the integration level of different products varies. Illustration by Xiang Xianzhi.

The new PC orders thus raise a bigger question: the extra money users spend on local AI, will finally go to the company that provides chips, the company that provides systems, or the company that sells the machine to users?

Both Apple and Xiaomi are trying to keep more links under their control. But the resources they already have, as well as the parts they need to spend extra money to supplement, are very different.

Apple controls the full stack system, Xiaomi reaches down to the underlying layer

Apple can adjust chips, systems and development tools at the same time, and it does not require users to make a new purchase of a dedicated AIPC to enable local model deployment.

Unified memory is Apple's killer feature: Apple's MLX framework allows CPU and GPU to use arrays in shared memory, reducing the need to transfer data between different computing units; MLX-LM provides tools for model operation, quantization and fine-tuning. Developers can download third-party open models and also use Apple's hardware-optimized tools to run them.

This path mainly uses CPU and GPU, not only relying on dedicated NPU. The Core AI framework previewed this year further optimizes unified memory and the Neural Engine. Apple can continuously supplement software around the same set of hardware, instead of re-looking for a new computing platform every time an AI feature is added.

The August update further differentiated its product tiers. The M6 Mac mini starts at 6999 yuan, with up to 32GB memory; the M5 Pro version supports up to 64GB. Users who need larger models and heavier development tasks can continue to choose Mac Studio with 128GB or 512GB configurations. The new product was open for pre-order on August 27, and is scheduled to be officially released on September 22.

Figure 3|The new Mac Studio offers M5 Max and M5 Ultra configurations. Image source: Apple Newsroom.

Users are stratified: ordinary users need more practical features in existing applications; developers care about whether they can deploy their own models and codes; research teams and professional creators are more likely to pay for large memory and continuous load performance. But Apple can cover users with different budgets with the same system, without rebuilding sales channels for each tier of users.

Mastering the full platform does not make Apple reject cooperation on external models. The third-generation foundational model announced in June was developed in cooperation with Google, including a 3B edge model and a 20B edge model that adopts a sparse structure and activates part of parameters according to requests.

Cooperation has changed the source of models, but the decision-making power over systems and products still rests with Apple. Which models are included in system services, what interfaces developers use to call them, and how applications get permission before reading files, are all arranged by its platform. External models can run on Mac, and users do not need to leave the Mac ecosystem because of that.

Therefore, Apple's commercial plan is not limited to selling memory on traditional Mac devices. It is also adding new usage scenarios to existing hardware, encouraging developers to continue building new features on its system. Models and frameworks are still gradually opening up with system versions, regions and testing progress, but the existing product lines and user relationships allow Apple to promote this work in batches.

Xiaomi wants to grasp similar initiative, but it is obviously taking a different path from Apple.

The division of labor of the three chips in AI Cube is very illustrative: O3 is for personal intelligent terminals, O100 is for edge AI acceleration, and D100 is for autonomous driving. Xiaomi has integrated several chip R&D lines into the same prototype to demonstrate collaborative computing capabilities. According to the information released in August, O3 has entered mass production, and O100 and D100 are planned to be commercially available next year. AI Cube itself is still an engineering prototype, with no price and release date announced.

Figure 4|The AI Cube engineering prototype. Image source: Xiaomi, reposted by IT Home.

The significance of self-development first lies in that Xiaomi can organize the design of a chip on its own, which does not mean that every component is built from scratch. The previous Xuanjie O1 used Arm's Cortex CPU, Immortalis GPU and CoreLink interconnect IP. For Xiaomi, investing in chip design is to arrange computing power, memory and power consumption for its own products on the basis of general technologies, and then be subject to the constraints of manufacturing, packaging and memory supply.

While the hardware side is still under development, Xiaomi has already started looking for applicable scenarios on Mijia for future tasks.

Miloco 2.0, released in June, connects the audio and video of Mijia cameras, the MiMo model and device control. It runs in the form of an OpenClaw plugin, trying to identify family members, remember their habits, and then control the devices at home. In the past, users needed to pre-set rules like "if X happens, turn on Y". Now Xiaomi is trying to let AI participate in understanding and arranging these things.

But there is an interesting point: Xiaomi's own installation guide recommends preparing a Mac mini that runs 24/7.

Figure 5|The official deployment document of Miloco 2.0 recommends that the device runs 24/7, and recommends Mac mini. Source: Xiaomi's official GitHub repository. Screenshot date: September 7, 2026.

Xiaomi can first develop the home AI software on other manufacturers' hardware. Miloco 2.0 currently mainly invokes cloud large models, while AI Cube demonstrates local computing capabilities. The two are not yet a fully integrated product. Xiaomi still needs to complete the whole path from software that can control home appliances to services that are worth buying a separate host to run.

As of the end of June 2026, the number of connected devices on Xiaomi's AIoT platform reached 1.1608 billion, excluding mobile phones and tablets; the number of users with 5 or more connected devices reached 24.6 million. These users are already using Mijia, so Xiaomi has the opportunity to deliver new services to them.

If AI begins to replace users to understand home status and arrange continuous tasks, the software responsible for making judgments will influence how the next device is selected and used. Xiaomi hopes to continue to control this entry point, rather than only providing controllable hardware for other manufacturers' agents.

Chip design adds the possibility for it to arrange computing according to its own scenarios. But Mijia users are not naturally buyers of AI hosts. AI Cube still needs to prove what work it can complete that is worth users paying extra for, compared with existing devices and cloud services. At present, the full collaboration mode between this prototype and HyperOS, Mijia and Miloco, as well as the support for third-party models and development tools, have not been fully launched.

Huawei has also integrated chips and systems into personal computers. The MateBook Pro is equipped with Kirin X90 Plus; the Xiaoyi local AI on HarmonyOS PC can offline complete part of document Q&A, summarization and rewriting. The "Local AI Model Management" feature of HarmonyOS 6 also allows users to download models in system settings, and then enable applications such as Chatbox and Cherry Studio to invoke them.

This allows system vendors to further participate in model distribution. Huawei can integrate local AI into the system operations familiar to office users, but current model management is still bounded by a single device, and cross-device sharing is not supported. The product has been launched, but there is still room for improvement in application coverage and cross-device experience.

Back to Apple and Xiaomi, the directions of the two companies are actually different here. Apple can first let the already sold Mac devices do more things, and then sell higher configurations to heavy users; Xiaomi already has home devices and users, but it still needs to prove that connecting dedicated computing hardware with home services is worth a new separate purchase.

Figure 6|The vertical integration routes of Apple and Xiaomi.

PC manufacturers begin to compete for AI entry points

Traditional PC manufacturers have come up with another approach: sharing upstream platforms, and then integrating deployment and services into the complete devices. Although the GB10 devices of different companies are under different brands, their chips and main development tools come from the same supply system.

Figure 7|Shared platforms and customized cooperation of PC manufacturers. Illustration by Xiang Xianzhi.

Sharing the platform saves the investment of developing chips and basic software from scratch, but also leaves part of the product decision-making power to the upstream. NVIDIA provides CUDA and model tools. Developers can verify models on the desktop first,