HomeArticle

From selling GPUs to selling Tokens, Fenghe Intelligence reconstructs the profit logic of AI Infra with a nine-layer architecture.

氪报2026-08-18 15:51
Fenghe Smart promotes token efficiency-focused operations, helping intelligent computing centers realize transformation and operational efficiency improvement.

As the demand for large model training and inference continues to grow, the construction of global artificial intelligence infrastructure is accelerating. From China to the Middle East, North America and Southeast Asia, massive GPU resources are being deployed to data centers, with trillions of funds pouring into AI infrastructure construction. The intelligent computing industry has also gradually moved from the early infrastructure construction stage to an operation stage that places greater emphasis on efficiency, stability and commercial returns.

In this process of change, the traditional "bare metal leasing" model that charges by GPU card hour, node or cabinet is facing new challenges. For intelligent computing centers, the core indicator to measure their operating capabilities is no longer just how many GPUs they own, but how many Tokens can be stably output per unit time, what the comprehensive cost per million Tokens is, and how much actual value each GPU can create.

Based on this industry trend, Fenghe Intelligent Technology (Shanghai) Co., Ltd. has proposed the concept of "Token Efficiency Operation Model". Through the nine-layer technology stack contained in AI Infra, the lean intelligent computing methodology and the authoritative test endorsement from the China Academy of Information and Communications Technology, it promotes the transformation of intelligent computing centers from "low-profit bare metal leasing" to "high-value Token factories".

Shift from computing power bare metal equipment delivery to efficient Token production capacity operation

Fenghe Intelligence has been operating in Zhangjiang Science City, Shanghai since 2021, focusing on AI Infra security and efficiency solutions. Its core team members have 20 years of working experience in international IT vendors, and have accumulated long-term experience in technical implementation, channel construction and large-scale project delivery.

"GPUs are the production equipment in artificial intelligence infrastructure, and Tokens are the digital products finally produced by the equipment. What customers really need is not the GPU equipment itself, but the model training results, inference service capabilities and stable and sustainable Token output." Based on this judgment, Fenghe Intelligence put forward the "lean intelligent computing" construction standard in 2023, hoping to reduce problems such as GPU idling, queuing, video memory fragments and inefficient operation through technical means such as resource scheduling, operating environment optimization, model deployment, inference acceleration and automated operation and maintenance, so as to improve the effective output capacity of intelligent computing centers.

"What intelligent computing centers need to solve in the future is not only the problem of resource construction, but also the problem of production capacity operation. For customers, what they need are training results and inference capabilities; for intelligent computing centers, what can really create commercial value is the Token finally output by GPUs." —— Fenghe Intelligence AI Engineering Center

Core Methodology: Build TEF, TFI and TFOM Operation System

In order to promote the standardization of operation indicators of intelligent computing centers, Fenghe Intelligence has built a Token factory operation methodology composed of TEF, TFI and TFOM, which integrates technical performance, service quality and financial returns into a unified evaluation system.

Among them, TEF, namely Token Efficiency Factor, mainly measures the proportion of the effective time that GPUs are actually used for Token production to the total running time. According to the company's project practice, the effective output ratio of GPUs in most intelligent computing centers is only 20% to 35% of the theoretical production capacity, and a large amount of computing power is consumed in links such as task waiting, scheduling gaps, video memory fragments and inefficient batch processing. Through lean scheduling and full-link optimization, the company puts forward the operation goal of increasing TEF from 30% to more than 65%.

TFI, namely Token Factory Index, further introduces service availability and Token sales rate on the basis of TEF, and measures the overall capacity commercialization level of intelligent computing centers through the formula of "TEF × Service Availability × Capacity Sales Rate".

TFOM, namely Token Factory Operation Model, further corresponds technical indicators such as TPS (Tokens Per Second), latency (first Token / single Token), TEF (Token Efficiency Factor), GPU utilization and TCO (Total Cost of Ownership) with financial indicators such as revenue, customer experience, gross profit margin, asset turnover rate and return on investment, helping investors, intelligent computing center operators and enterprise managers to intuitively evaluate the project operation quality.

Taking a cluster composed of 1000 B200 GPUs as an example, according to the company's calculation method, if TEF is increased from 30% to 65%, the increment of its effective output is equivalent to adding about 1167 GPUs — no need for procurement, no waiting for goods, no additional computer room construction, directly saving hundreds of millions of yuan in capital expenditure (CapEx).

Technical Moat: Nine-layer Full-stack Self-developed Architecture Opens Up the Technical Link of Intelligent Computing Centers

At the technical level, Fenghe Intelligence's lean intelligent computing products cover the nine-layer technology stack of the full life cycle of artificial intelligence infrastructure.

Starting from the physical infrastructure layer, this architecture sequentially covers the hardware driver layer, system runtime layer, resource scheduling layer, storage data layer, training framework layer, model management layer, inference service layer and application service layer, forming a complete technical link from computer room cooling, power supply management, GPU driver adaptation to model deployment, inference engine, API gateway and multi-modal scheduling service delivery.

Different from technical solutions that only focus on resource management, model training or inference services, the nine-layer architecture emphasizes the collaborative interaction between all layers. Upper-layer scheduling can perceive the underlying hardware status, underlying faults can be automatically resolved by the upper layer, the KV-Cache strategy of the inference layer is linked with the data prefetching of the storage layer, the gradient aggregation of the training framework is connected with the NCCL optimization of the communication layer, and there is no breakpoint in the entire link.

At the same time, relying on the Ti-X scheduling engine, ADS automatic deployment system, KV-Cache optimization engine and full-link automated operation and maintenance capabilities, Fenghe Intelligence has formed a complete service system from intelligent computing center design, deployment, optimization to continuous operation, and has the technical capability to manage more than 10,000 GPUs in a single cluster.

Authoritative Endorsement: AISHPerf Test from CAICT Verifies System Service Capability

In May 2026, Fenghe Intelligence accepted relevant tests from the China Academy of Information and Communications Technology (CAICT) Thiel Laboratory. The tests were carried out based on the AISHPerf artificial intelligence system performance test platform, covering multiple dimensions such as software and hardware configuration verification, software framework functions, client service capabilities, monitoring dashboard functions and large model inference performance.

In the large model inference performance test, the combination of Fenghe Tiansui X500G2 Ti-X inference architecture and S300G2 high-performance storage was used to verify the ultra-long context scenario. The actual measurement shows that the maximum reduction of first Token and end-to-end latency exceeds 80%, and the Token output throughput is increased by 96%-140%.

According to the company's operating data in actual customer projects, after optimization with the nine-layer architecture, the throughput of the relevant inference system is increased by more than 200%, and the first Token latency is reduced by about 80%. In training and cluster operation scenarios, the overall resource utilization rate is increased by more than 30%, the hardware failure rate is reduced by 50%, and the manual operation and maintenance cost is reduced by 70%.

These test and project data provide quantitative reference for the continuous operation, resource scheduling, model training and high-concurrency inference services of large-scale intelligent computing centers, and also provide technical capability support for the company to participate in the construction of intelligent computing projects of large enterprises, governments and central state-owned enterprises.

Build a Service System Covering Domestic and Overseas Markets

Fenghe Intelligence mainly provides full-life cycle services for owners of large-scale intelligent computing centers, including early-stage TCO calculation, architecture planning and equipment selection, mid-stage nine-layer technology stack deployment, cluster joint debugging and performance optimization, and later-stage SLA guarantee, efficiency monitoring and continuous operation optimization.

In May 2026, the company set up its international business export and Asia-Pacific service center in Yazhou Bay National Science and Technology City, Sanya, Hainan, and plans to export lean intelligent computing capabilities to markets with rapidly growing demand for artificial intelligence infrastructure such as Southeast Asia, the Middle East and North America. In the same period, Fenghe Intelligence signed an overseas AIDC strategic cooperation with Unisplendour Intelligent Computing, and the two sides will combine global channel resources, intelligent computing infrastructure construction capabilities and full-stack operation technologies to jointly expand the overseas artificial intelligence data center market.

As the AI Infra industry moves from large-scale construction to refined operation stage, the competition focus of intelligent computing centers is shifting from the number of GPUs to the actual output capacity per unit of computing power. How to increase Token output, reduce the cost per unit Token, ensure service stability, and form sustainable commercial returns will become an important topic in the next stage of the industry.

Fenghe Intelligence said that in the future, it will continue to improve the technical exploration of the nine-layer technology stack and the Token factory operation efficiency model, promote the seamless integration of technical indicators and operation indicators, let each GPU release the ultimate Token output, and upgrade the intelligent computing center from a computing power resource center to a sustainable Token digital production capacity platform.