AI on China-market iPhones is just around the corner, revealing how Apple develops its "China-exclusive" model.
As the countdown to the new iPhone launch begins, Chinese users are once more concerned about this question:
Will the China-region version of Apple Intelligence be usable at all?
According to an exclusive Reuters report on August 14, Apple has trained a large language model tailored for the Chinese market. Three people familiar with the matter revealed that this model is co-developed by Apple and Alibaba, and Alibaba also provided support for the model training.
Neither Apple nor Alibaba has responded to requests for comment. As the countdown to the new iPhone launch begins, Chinese users are once more concerned about this question: Will the China-region version of Apple Intelligence be usable at all?
This may differ from the public's speculation about the China-region AI service. The previously disclosed solution shows that Apple mainly relies on Chinese partners to provide model capabilities. Alibaba's Qwen will be integrated into the Chinese version of Apple Intelligence, and Baidu's technology will also be involved in some functions.
Now, Apple's own China-exclusive model has also been added to this system. It turns out that the China-region version of Apple Intelligence may be supported by technologies from Apple's self-developed model, Qwen, and Baidu at the same time.
What does Apple's own model mean?
"Apple's own model" is easily interpreted as a self-developed model, but in fact it does not mean that all training work is completed independently by Apple, and the model does not necessarily start training from random parameters.
At this year's WWDC, Apple officially released the new generation of Apple Intelligence, whose model was jointly trained by Apple and Google based on the Gemini model and Google Cloud technology, rather than directly embedding Gemini into iOS as many people previously assumed.
Apple named this model the third-generation Apple Foundation Models, or AFM 3 for short.
Apple is responsible for defining the model architecture, product capabilities, training process, privacy mechanism, and software and hardware adaptation, while Google provides model technology, cloud computing power and part of the infrastructure.
It is very likely that Alibaba's provision of training support for Apple's China-exclusive model follows the same cooperation pattern.
Alibaba can cover many links in this process, including Chinese corpus processing, training infrastructure, local knowledge enhancement, safety alignment, model evaluation and synthetic data generation, all of which may require Alibaba's technology and resources.
Apple may also use mature models to generate training data, and then migrate part of the capabilities to small models suitable for end-side operation. However, Reuters did not disclose the source of the base weights of the China-specific model, nor did it confirm whether it is developed based on Qwen.
Only two points can be confirmed at present: Apple has an exclusive model for the Chinese market, and Alibaba participated in its development and training.
The technical direction of Apple's model training in China has long been hinted at
Referring to the AFM 3 technical materials publicly released by Apple this year, we can roughly infer what the model trained by Apple in China will look like.
Apple AFM 3 is a family of 5 models in total.
AFM 3 Core is an end-side dense model with 3 billion parameters, which mainly handles latency-sensitive tasks that can be completed locally.
AFM 3 Core Advanced has a total of 20 billion parameters and adopts a sparse activation architecture. Only 1 billion to 4 billion parameters are activated for each request.
Traditional Mixture-of-Experts models usually reselect experts when generating each token. This method works well on cloud servers, but mobile phones are limited by memory capacity and data transmission speed.
AFM 3 Core Advanced uses the Instruction-Following Pruning technology proposed by Apple.
The full model weights are stored in NAND flash memory. After receiving an instruction, the model first selects the expert weights that need to be run according to the entire request, and then loads the relevant parts into DRAM.
During the generation process, the model reselects experts only when necessary.
There is also a group of shared experts in the architecture that always participate in the calculation, responsible for stably processing general capabilities. Other routing experts are dynamically loaded according to the task type.
In this way, the device can retain the capability scope of 20 billion parameters, while keeping the computation amount and memory usage of a single request at a relatively low level.
Notification summarization, text rewriting and simple Q&A can use fewer activated parameters. Complex context processing, voice processing and cross-app requests can call more parameters, and the same model will adjust its actual operation scale according to the task.
This architecture has strong practical significance for the China-exclusive model.
Apple needs to improve Chinese understanding and generation capabilities without significantly increasing the model size, memory footprint and power consumption. If the China-specific model follows a similar design, Apple can leave more Chinese processing tasks to be completed on the device side.
In addition, the AFM 3 family includes 3 more server-side models.
AFM 3 Cloud is responsible for regular cloud tasks. It adopts an improved Parallel-Track Mixture-of-Experts architecture, with a focus on enhancing multi-modal understanding, long-context memory and complex request processing.
ADM 3 Cloud is dedicated to image generation, photo editing and Genmoji.
AFM 3 Cloud Pro is oriented to complex reasoning and Agent tool invocation, and it is also the most capable model in the AFM 3 family.
Cloud Pro runs on NVIDIA GPUs on Google Cloud, and is also incorporated into Apple's Private Cloud Compute system.
Apple claims that when requests enter Private Cloud Compute, user data will not be saved, nor will it be accessible to Apple or other institutions.
However, there is no public information about where the Chinese version of the cloud model will be deployed and whether it will fully adopt the same Private Cloud Compute architecture.
This part will directly affect the privacy statement, response speed and available functions of the China-region Apple Intelligence.
Pay attention to Apple's training process: different models of AFM 3 will first share a general foundation, and then be trained separately for end-side, cloud-side, voice, image and complex reasoning scenarios.
After pre-training is completed, the models will also undergo supervised fine-tuning, multi-stage reinforcement learning and quantization-aware training.
The last step is hardware optimization. The end-side models are adjusted for Apple chips, while Cloud Pro is optimized for NVIDIA GPUs.
The public training data disclosed by Apple includes public data, licensed or purchased data, open source data, special research data and synthetic data. Apple also explicitly states that users' private personal data and user interaction records will not be used to train the foundation models.
Although the division of labor of these Apple models is not yet clear, APPSO speculates that the China-region Apple Intelligence is more likely to adopt a multi-model scheduling mechanism.
Tasks such as text summarization, writing assistance and screen understanding can be preferentially processed by the end-side model. Requests that require more computing power may be sent to the cloud model controlled by Apple.
When external knowledge or specific generation capabilities are involved, the system can also call Qwen. Users do not need to understand the difference between the back-end models, and the system will select the processing path according to the task content, device status, network environment and privacy requirements.
This division of labor is consistent with Apple's existing end-cloud collaboration framework, but the specific implementation remains to be confirmed by Apple.
It is particularly worth noting that Apple briefly released a Chinese support guide last week.
The guide introduced how Mac users in mainland China can connect Qwen to Siri and writing tools. Subsequently, the page was deleted, and Apple did not explain the reason.
This document shows that Qwen may be accessed as a clear external model entry in Apple's system.
Whether it will also participate in built-in functions, whether separate authorization is required, and how it switches with Apple's exclusive model, may not be known until the service is officially launched.
When buying an iPhone of the same generation, are you still using the same product?
In July this year, the Cyberspace Administration of China announced the filing information of 7 new mobile end-side generative AI services.
"Apple Intelligence" appears on the list. The official filing announcement confirms that Apple has passed the most critical compliance procedure for the Chinese market.
The announcement did not provide the official launch date, nor did it disclose the corresponding iOS version and supported device list.
According to the information obtained by Reuters, Apple Intelligence is expected to enter the Chinese market with an iOS system update in the next few months.
Apple has previously announced that iOS 27 and the new generation of Apple Intelligence will be launched this fall, and Siri AI will be opened as a beta version later this year.
Apple still stated during WWDC 26 that these new functions will not enter the Chinese market immediately.
The probability that the China-region Apple Intelligence will be launched simultaneously with the new iPhone seems not high. Apple may also release the hardware first, and then open the relevant functions through a subsequent iOS 27 update.
At present, the global version of Apple Intelligence supports iPhone 15 Pro, iPhone 15 Pro Max, iPhone 16 series and later models. The supported device list of the Chinese version is also unknown.
Whether existing devices can get full functions in the first batch may depend on the requirements of the end-side model for memory, storage space and chip computing power.
Different functions may also be launched in batches. The fact that the end-side writing tool has completed the filing does not mean that the new Siri AI, complex Agent capabilities and all cloud functions will be available on the same day.
In the past, the differences of iPhones in different regions were mainly reflected in a few hardware specifications and network services.
But the advent of AI has changed this situation. Two iPhones with the same chip and running the same iOS version may have completely different system capabilities due to different models, cloud facilities and data processing rules.
Apple's separate training of models for China also means that Apple Intelligence is evolving from a globally unified service to a regionalized architecture that allows replacement of models and infrastructure.
This will give Apple higher control, but also increase the complexity of system maintenance.
For example, can Apple's models for overseas and China markets get new capabilities synchronously? Every time Siri AI adds new functions, does Apple need to complete local adaptation and re-evaluation?
These issues will directly affect the rhythm of subsequent iOS updates.
Therefore, the launch time of Apple's China-region AI is not only related to when Chinese users can use Siri AI, but also to a more long-lasting issue that affects user experience:
When software experience is increasingly dependent on models and cloud services, buying an iPhone of the same generation no longer necessarily means using the exact same product.
This article is from WeChat Official Account "APPSO", author: Discovering Tomorrow's Products, published by 36Kr with authorization.