HomeArticle

From selling GPUs to selling Tokens, Fenghe Intelligence reconstructs the profit logic of AI Infra with a nine-layer architecture.

氪报2026-08-11 14:35
Fenghe Smart Introduces Token Efficiency Model to Empower Intelligent Computing Centers to Transform to High-Value Operations

With the growing demand for large model training and inference, the construction of global artificial intelligence infrastructure is accelerating. From China to the Middle East, North America and Southeast Asia, a large number of GPU resources are being deployed in data centers, trillions of funds are pouring into AI infrastructure construction, and the intelligent computing industry has gradually moved from the early stage of infrastructure construction to an operation stage that pays more attention to efficiency, stability and commercial returns.

In this process, the traditional "bare metal leasing" model that charges by GPU card hour, node or cabinet is facing new challenges. For intelligent computing centers, the core indicator to measure their operating capabilities is no longer just how many GPUs they own, but how many Tokens can be stably output per unit time, what the comprehensive cost per million Tokens is, and what actual value each GPU can create.

Based on this industry trend, Fenghe Intelligent Technology (Shanghai) Co., Ltd. has put forward the concept of "Token Efficiency Operation Model". With the nine-layer technology stack included in AI Infra, the lean intelligent computing methodology and the authoritative test endorsement from the China Academy of Information and Communications Technology, it promotes the transformation of intelligent computing centers from "low-profit bare metal leasing" to "high-value Token factories".

Shift from computing power bare metal equipment delivery to efficient Token production capacity operation

Since its operation in Zhangjiang Science City, Shanghai in 2021, Fenghe Intelligence has focused on AI Infra security and efficiency solutions. Its core team members have 20 years of working experience in international IT vendors, and have accumulated long-term experience in technical implementation, channel construction and large-scale project delivery.

Fenghe Intelligence believes that GPU is the production equipment in artificial intelligence infrastructure, and Token is the digital product finally produced by the equipment. What customers really need is not the GPU equipment itself, but the model training results, inference service capabilities and stable and sustainable Token output.

Based on this judgment, Fenghe Intelligence proposed the "lean intelligent computing" construction standard in 2023, hoping to reduce problems such as GPU idling, queuing, video memory fragments and inefficient operation through technical means such as resource scheduling, operating environment optimization, model deployment, inference acceleration and automated operation and maintenance, so as to improve the effective output capacity of intelligent computing centers.

"What intelligent computing centers need to solve in the future is not only the problem of resource construction, but also the problem of production capacity operation. For customers, what they need are training results and inference capabilities; for intelligent computing centers, what can really create commercial value is the Token finally output by GPUs." -- AI Engineering Center of Fenghe Intelligence

Core Methodology: Build TEF, TFI and TFOM Operation System

In order to promote the standardization of operation indicators of intelligent computing centers, Fenghe Intelligence has built a Token factory operation methodology consisting of TEF, TFI and TFOM, which brings technical performance, service quality and financial returns into a unified evaluation system.

Among them, TEF, namely Token Efficiency Factor, mainly measures the proportion of the effective time that GPU actually uses for Token production in the total running time. According to the company's project practice, the effective output ratio of GPUs in most intelligent computing centers is only 20% to 35% of the theoretical production capacity, and a large amount of computing power is consumed in links such as task waiting, scheduling gaps, video memory fragments and inefficient batch processing. Through lean scheduling and full-link optimization, the company puts forward the operation goal of increasing TEF from 30% to more than 65%.

TFI, namely Token Factory Index, further introduces service availability and Token sales rate on the basis of TEF, and measures the overall commercialization level of production capacity of intelligent computing centers through the formula of "TEF × Service Availability × Production Capacity Sales Rate".

TFOM, namely Token Factory Operation Model, further maps technical indicators such as TPS (Tokens per Second), latency (first Token / single Token), TEF (Token Efficiency Factor), GPU utilization and TCO (Total Cost of Ownership) to financial indicators such as revenue, customer experience, gross profit margin, asset turnover and return on investment, helping investors, intelligent computing center operators and enterprise managers to evaluate the project operation quality more intuitively.

Take a cluster composed of 1000 B200 GPUs as an example. According to the company's calculation method, if TEF is increased from 30% to 65%, the effective output increment is equivalent to adding about 1167 GPUs -- no need for procurement, no waiting for goods, no additional computer room construction, directly saving hundreds of millions of yuan in capital expenditure (CapEx).

Technical Moat: Nine-layer Full-stack Self-developed Architecture Opens Up the Technical Link of Intelligent Computing Centers

At the technical level, Fenghe Intelligence's lean intelligent computing products cover the nine-layer technology stack of the full life cycle of artificial intelligence infrastructure.

Starting from the physical infrastructure layer, this architecture successively covers the hardware driver layer, system runtime layer, resource scheduling layer, storage data layer, training framework layer, model management layer, inference service layer and application service layer, forming a complete technical link from computer room cooling, power supply management, GPU driver adaptation to model deployment, inference engine, API gateway and multi-modal scheduling service delivery.

Different from technical solutions that only focus on resource management, model training or inference services, the nine-layer architecture emphasizes the collaboration and linkage between all layers. The upper-layer scheduling can perceive the status of the underlying hardware, the underlying faults can be automatically resolved by the upper layer, the KV-Cache strategy of the inference layer is linked with the data prefetching of the storage layer, and the gradient aggregation of the training framework is connected with the NCCL optimization of the communication layer. There is no breakpoint in the whole link. For example, the resource scheduling system can perceive the status of the underlying hardware and automatically migrate tasks according to the fault situation; the KV-Cache optimization of the inference layer can cooperate with the data prefetching mechanism of the storage layer; the gradient aggregation capability in the training framework can also be linked with the underlying communication optimization mechanism.

Relying on the Ti-X scheduling engine, ADS automatic deployment system, KV-Cache optimization engine and full-link automated operation and maintenance capabilities, Fenghe Intelligence has formed a complete service system from intelligent computing center design, deployment, tuning to continuous operation, and has the technical capability to manage more than 10,000 GPUs in a single cluster.

Authoritative Endorsement: AISPerf Test from China Academy of Information and Communications Technology Verifies System Service Capability

In May 2026, Fenghe Intelligence received relevant tests from the Cloud Computing and Big Data Research Institute of the China Academy of Information and Communications Technology. The tests were carried out based on the AISPerf artificial intelligence system performance test platform, covering multiple dimensions such as hardware and software configuration verification, software framework functions, client service capabilities, monitoring dashboard functions and large model inference performance.

In the large model inference performance test, the Qwen/DeepSeek-R1-Distill-Qwen-32B model was adopted, and 14,400 requests were completed under the full-load concurrency scenario -- 0 failures. The test data shows that the system still maintains high-throughput output under full load, the average latency of the first Token is controlled in a reasonable range, and the generation latency of a single Token is only about 12 milliseconds.

According to the operation data of the company's actual customer projects, after the optimization of the nine-layer architecture, the throughput of the relevant inference system has increased by more than 200%, and the first Token latency has been reduced by about 40%. In training and cluster operation scenarios, the overall resource utilization rate has increased by more than 30%, the hardware failure rate has been reduced by 50%, and the manual operation and maintenance cost has been reduced by 70%.

These test and project data provide quantitative references for the continuous operation, resource scheduling, model training and high-concurrency inference services of large-scale intelligent computing centers, and also provide technical capability support for the company to participate in the construction of intelligent computing projects of large enterprises, governments and central state-owned enterprises.

Build a Service System Covering Domestic and Overseas Markets

Fenghe Intelligence mainly provides full-life cycle services for the owners of large intelligent computing centers, including early-stage TCO calculation, architecture planning and equipment selection, mid-stage nine-layer technology stack deployment, cluster joint debugging and performance optimization, and later-stage SLA guarantee, efficiency monitoring and continuous operation optimization.

In May 2026, the company's international business export and Asia-Pacific service center was settled in Yazhouwan National Science and Technology City, Sanya, Hainan, planning to export lean intelligent computing capabilities to markets with rapidly growing demand for artificial intelligence infrastructure such as Southeast Asia, the Middle East and North America.

In the same period, Fenghe Intelligence signed an overseas AIDC strategic cooperation with Unigroup Intelligent Computing. The two parties will combine global channel resources, intelligent computing infrastructure construction capabilities and full-stack operation technologies to jointly expand the overseas artificial intelligence data center market.

As the AI Infra industry moves from large-scale construction to refined operation, the competition focus of intelligent computing centers is shifting from the number of GPUs to the actual output capacity per unit of computing power. How to increase Token output, reduce the cost per unit Token, ensure service stability, and form sustainable commercial returns will become an important topic for the next stage of the industry.

Fenghe Intelligence said that in the future, it will continue to improve the technical exploration of the nine-layer technology stack and the Token factory operation efficiency model, promote the seamless integration of technical indicators and operation indicators, let each GPU release the ultimate Token output, and upgrade the intelligent computing center from a computing power resource center to a sustainable Token digital production capacity platform.