HomeArticle

Apple's real AI trump card is not Siri at all.

爱范儿2026-09-28 11:55
Apple, which withdrew from the server market 15 years ago, is poised to make a comeback driven by AI.

Is Apple Falling Behind in the AI Era?

This question is now posed to Tim Cook's successor John Ternus. The repeated delays of AI Siri in the past two years have made many people pessimistic about the answer, while at the same time, Mac mini has become synonymous with AI PC and is almost impossible to get due to high demand.

Recently, veteran tech analyst Ben Thompson put forward a provocative argument on his podcast: Apple does not need to develop AI on its own at all.

He believes that Apple, which controls a huge user entry point, only needs to procure AI capabilities and focus on what it is better at. The biggest risk is that if future AI is no longer centered on mobile phones, Apple may be subverted.

At this month's Apple launch event, Ternus redefined iPhone as the Intelligent Personal Hub, aiming to connect a series of upcoming AI hardware and resist the risk of AI decentralization.

It's not that Apple doesn't need to work on AI, but its route will most likely not be competing for model development, and will very likely return to the hardware itself.

Not long ago, the news that Apple was exposed to return to the enterprise server market is a signal. Apple left the server market 15 years ago, and now it is coming back because of AI.

This device, which may appear as early as 2029, will be equipped with two or four M8 Ultra chips, and may even use NVIDIA's NVLink Fusion to connect multiple chips into a larger system.

Contrary to the widespread rumor that it is used for self-developed models, what Apple wants to do is to run pre-trained AI models for enterprises. The project is still in the embryonic stage of discussion, but has received direct support from new CEO John Ternus.

The last time Apple sold a dedicated server was the Xserve, which was discontinued in 2011. Fifteen years later, why did it choose to re-enter the computer room in the AI era?

It seems that the answer is related to the booming server business, but when it comes to Apple, it is obviously not that simple.

Local AI Is Still a Geeky Practice

Since the beginning of this year, thanks to the "lobster craze" in early 2026, Mac mini has unexpectedly become a popular device for personal AI players. After its launch, OpenClaw has received more than 100,000 GitHub Stars and attracted 2 million visitors in a single week, which shows that users do want to have a private Agent that runs continuously and retains memories.

Some people buy high-memory versions to run open source models at home; some set it as an all-day online Agent to keep chat records, files and long-term memories on their own machines. According to the calculation of The Wall Street Journal, Mac mini previously only accounted for about 3% of Apple's US Mac sales, but the demand brought by resident agents such as OpenClaw has led to obvious shortages of some high-memory versions, and prices have also risen.

For a while, Mac mini became a rare and valuable commodity, and people who couldn't buy it had to find other solutions, such as renting a VPS to put the Agent and memories on a remote server. The main customers of VPS were originally developers and enterprises. The whole process of renting servers, connecting remote terminals, deploying software, and managing ports and keys is relatively complicated, which requires a large learning cost for individual users.

But AI is pushing a group of "newbie" users into this quagmire at great security risk. In early 2026, security researchers found about 175,000 misconfigured Ollama services on the public network, many of which came from home networks, VPS and cloud hosts; they were originally supposed to respond only locally, but were directly exposed to the Internet due to user setting errors. TechRadar

It can be said that VPS is a quite geeky way of playing, and is not suitable for the default entry of consumer-grade AI, which is why cloud AI has always been the mainstream of the market. People don't need to judge where the model should be installed, nor do they need to consider memory, bandwidth and electricity costs. They can start a conversation by opening ChatGPT, Claude or Doubao. The larger the model and the more complex the task, the more obvious the advantages of the cloud.

However, personal local AI has formed a real demand, and people do begin to need an AI that is "held in the palm of their hand". Just from the shortage of Mac mini to 170,000 Ollama services running unprotected on the public network, the result of leaving this demand to users to handle on their own is a mix of enthusiasm and risks.

This problem is not only for individual users, enterprises have long encountered its amplified version, but enterprises have the budget to bear it.

Apple to Deliver the AI That Enterprises Want

For compliance and data security reasons, many enterprises are very cautious about handing over inference tasks to public clouds; but if they want to keep tasks in their own computer rooms, they have to select chips by themselves, build the software stack by themselves, and maintain the inference environment by themselves. At present, there is almost only one way to build from scratch around NVIDIA's ecosystem, and no one sells an out-of-the-box local inference all-in-one machine.

So what Apple actually wants to sell is exactly this. The device mentioned in the report runs pre-trained models for enterprises — note that it is for inference, not training.

The difference between the two lies not only in computing power, but also in where the data goes. Training must gather corpus into a super-large cluster, where Apple has no advantages at all, but it has a ready-made solid foundation in inference: the unified memory architecture it has adhered to since the M1. This allows the CPU and GPU to share the same memory, and a single Mac Studio can hold model weights that originally required multiple professional graphics cards to load.

Take the latest Mac Studio as an example. The M5 Ultra version can be configured with up to 512GB of unified memory, with a memory bandwidth of 1.2TB/s, and supports multiple units to form a cluster via Thunderbolt 5. In contrast, even for top professional graphics cards, the video memory of a single card is usually only in the order of tens of GB. To load a model of the same scale, you have to split it into multiple cards first and then find a way to combine them back. Apple's official statement also directly describes the new Mac Studio as the ultimate desktop for edge-side AI.

The unified memory, which was originally a design choice for desktop creation, has shown new value in the era of large model operation. Therefore, Apple is not trying to compete with AWS for the cloud business. It sells the machine itself, not computing power. The machine is placed in the enterprise's own computer room and runs the enterprise's own data. From this perspective, it is more like the return of Xserve, except that its purpose is changed to AI inference.

The Xserve server that Apple once developed

For individual users, you don't have to learn to deploy models by yourself for local AI; the same is true for enterprise users, who don't have to build a set of AI infrastructure by themselves — Apple can provide a hardware and software integrated solution.

The New CEO Has Made the Decision

Apple's move to return to the server market is clearly supported by the new CEO.

This year, John Ternus succeeded Tim Cook to take charge of Apple. He is a mechanical engineer by background, who has long been in charge of hardware products such as iPhone, iPad, Mac, Apple Watch and AirPods. He is the first Apple helmsman to rise to the CEO position from the hardware system in the past 30 years.

The just-concluded fall launch event was his first public appearance, where iPhone Duo, iPhone 18 Pro & Pro Max, Apple Watch Series 12 & Ultra 4, and AirPods 5 were all unveiled on the stage.

This series of products all reflect a strong feature of "making decisions for users". Take the A20 Pro chip as an example, there are many trade-offs that users cannot see: the memory bandwidth is increased by 50%, the NPU is made into 32 cores with twice the performance of the A18 Pro, and the area of the new generation of vapor chamber is directly three times that of the previous generation. These hardware-level improvements potentially allow more models to run on this phone.

Apple didn't ask users if they were willing to pay for edge-side inference, it directly welded the answer in its mind into the chip.

"Making decisions for users" is a good summary of this year's product style. Different from Tim Cook, it reminds people of a past figure: Steve Jobs.

The first-generation iMac directly removed the floppy drive, when almost all files were transferred through it; the first-generation iPhone canceled the physical keyboard and stylus, which were the most proud features of business phones at that time; MacBook Air removed the optical drive, and iPad and iPhone have always rejected Flash...

This is a temperament that Apple has not shown for a long time: It does not base its decisions on the mainstream preferences of the market, but first judges what it thinks is a good way of use, and then directly presents it to users.

However, the reason why John Ternus can be so resolute is the confidence Apple gives him, as well as the mutual need between him and the company. When Ternus took office this spring, Apple's AI was in an awkward period: the new Siri was delayed again and again, and both self-developed and joint models were reported to encounter difficulties. Earlier this year, Apple signed an agreement with Google for about one billion US dollars per year to use Gemini to power Siri, but no one expected that Gemini would underperform significantly in the second half of the year...

It was not until the new Siri AI was unveiled at WWDC 2026 that the situation was barely stabilized. Of course, the evaluation from Mark Gurman of Bloomberg is only "just about usable".

In other words, at the model level, Apple has admitted that it needs external support. What it can still fully control is the hardware, and where the computing determined by the hardware takes place. For a CEO with a hardware background, betting on this layer is a reasonable and possibly the only choice.

This is a route that only Apple