HomeArticle

Four paths for domestic GPUs

王智远2026-09-29 09:35
Four paths, three hurdles and one measuring ruler

The domestic GPU sector sounds like a competition, right?

If you take a closer look, it is not the case at all. This is four completely different development paths. First, let's look at the basic context: NVIDIA's market share in China's AI chip market was still around 95% in 2022, and by mid-2026, Bernstein's forecast shows that this figure has dropped to 8% overall, with the high-end market almost fully replaced by domestic products.

From another statistical caliber, Reuters estimates that NVIDIA still accounts for about 55% of the overall market. Combining the two sets of data, we can see that the high-end market is dominated by domestic players, while the mid-to-low-end market is still in a fierce melee.

The four leading domestic GPU vendors also take different technical routes.

Moore Threads follows the full-featured GPU route, covering graphics rendering, AI training and inference, and aims to build a complete system from chip design to software ecosystem, taking after NVIDIA's development path.

It has a high ceiling, but covering all scenarios means that it has to build the chip architecture, software stack and developer ecosystem from scratch. The moat that NVIDIA has spent more than 20 years building needs to be fully replicated from the ground up.

MetaX, on the other hand, chooses the CUDA-compatible route.

The code that originally ran on NVIDIA's products can be migrated to MetaX's cards with the lowest possible migration cost. The trade-off is that it will always be catching up with NVIDIA's iteration pace: every time NVIDIA launches a new generation of products, MetaX has to follow up with its corresponding update.

Its self-developed software stack MXMACA has achieved four layers of adaptation, covering drivers, compilers, operators and frameworks, and has been compatible with more than 40 AI frameworks and over 1000 models.

To put it bluntly, MetaX's core selling point is "no code modification required".

Biren has the most distinctive strategy: it bets on chiplet technology. Since a single chip cannot be made infinitely large, it combines multiple small chips together like building blocks.

Combined with its self-developed NPO optical interconnection technology, the optical engine is directly integrated into the GPU, bypassing the traditional DSP chip to achieve low-latency and long-distance transmission. It launched a 1024-card supernode solution at this year's WAIC (World Artificial Intelligence Conference).

Its bet is on large-scale training scenarios: if a single card cannot provide sufficient performance, a cluster can make up for the gap. However, while assembling chips is relatively easy, the inter-chip communication delay and consistency are the real pain points, and NPO is designed to solve this problem.

Enflame takes the most unique path: it develops a self-designed DSA architecture, which is specially optimized for inference scenarios and is not compatible with CUDA, so users need to rewrite their code, but it delivers much higher efficiency in inference scenarios.

Trading general-purpose performance for scenario-specific efficiency, this trade-off itself shows its firm confidence in the market judgment. It focuses on the inference track, and is deeply bound to its major client Tencent. According to the data in its prospectus, 83.79% of its revenue comes from Tencent. This deep partnership makes Tencent its largest test and application field.

All four companies are called "domestic GPU vendors", but their end points are completely different: Moore Threads aims to become the NVIDIA of China, MetaX focuses on the lowest code migration cost, Biren bets on large-scale training scenarios, and Enflame is dedicated to the inference track.

As you can see, the four paths lead to four different destinations, which is by no means the same competition.

How are these destinations determined? They are decided by the market structure. In 2025, inference scenarios have accounted for 67% of the AI server market, while training only accounts for 33%. The market size of inference demand is 4 to 5 times that of training.

Enflame's specialization in inference is exactly the fastest-growing segment. Huawei and Cambricon take the full-stack route covering both training and inference, so they have a foothold in both directions. In any case, no matter how different the four paths are, there is a common reality that cannot be avoided. Let's look at a set of data:

In 2025, the shipment of domestic AI accelerators reached 1.65 million units. Huawei Ascend ranked first with 810,000 units, T-Head ranked second with 265,000 units, and Cambricon and Kunlunxin tied for third with 116,000 units each. The total shipment of the four "small leading GPU vendors" is less than 7% of the overall market.

7% — just think about it, Huawei alone accounts for half of the market share in this sector. The real protagonist of this track is Huawei, and the four small leading vendors are more like four different tributaries with very limited scale.

As for NVIDIA, from the perspective of the overall market, it still accounts for about 55% of the share, but the high-end market has been almost fully replaced by domestic products. Combining the two sets of data, we can see that the high-end market has almost become a separate market, while the mid-to-low-end market is still in a fierce, immature competition.

Interestingly, no matter which path they take, the bottlenecks they encounter are exactly the same.

......

The first bottleneck is manufacturing process. All domestic GPU vendors need to find foundries for production, and SMIC (Semiconductor Manufacturing International Corporation) is the unavoidable option.

The production capacity of advanced manufacturing processes is tight, and Huawei alone has locked in about 43% of it. For every 10 wafers produced by SMIC with advanced process, more than 4 are for Huawei, and the rest are queued for other GPU companies.

This problem cannot be solved overnight. Building a wafer fab requires tens of billions of RMB in investment and takes 3 to 5 years, which cannot be accelerated simply by financing.

According to the CHIPS IV report released by Goldman Sachs in September 2026, by 2035, China's gap in advanced manufacturing processes will drop from 92% to 34%. The trend is positive, but for now, the four vendors are still waiting in the production queue.

Moore Threads' inventory turnover days have reached 658 days, with nearly two years of production capacity piled up in the warehouse, occupying a huge amount of cash flow. Stockpiling as much as possible once they get production capacity is a survival strategy forced by the tight supply situation.

The second bottleneck is the ecosystem. Every company has its own software stack, which is not compatible with others.

If an AI company wants to use domestic GPUs, it has to use the toolchain of the vendor it chooses, rewrite its code and re-adapt the frameworks. If it switches to another supplier, it has to go through a whole new round of adaptation and debugging. The learning cost of the team and the time spent on troubleshooting are all hidden expenses.

When making procurement decisions, enterprises cannot only look at the parameters of a single card, but also need to evaluate whether the entire toolchain is easy to use and whether the team can get familiar with it quickly.

There is a very accurate saying: the gap in hardware is shrinking, and the real bottleneck lies in the software ecosystem. There is a huge gap between "usable" and "user-friendly".

Just think about it, four companies with four independent software stacks means that each of them is independently building an incomplete CUDA from scratch.

NVIDIA's moat is not only the toolchain itself, but also the operator libraries, debugging experience and community Q&A content accumulated by millions of developers around the world over more than 20 years. There is no shortcut to build these assets.

All the pitfalls an engineer has stepped on in the NVIDIA ecosystem and the muscle memory he has accumulated will be reset to zero when he switches to a domestic GPU.

Moreover, every company is reinventing the wheel, developing compilers, operator libraries and debugging tools from scratch. Resources are scattered across four parallel tracks. Originally, concentrating 10 to 20 billion RMB of investment could build a usable ecosystem, but now the total investment is divided into four parts, with only 2 to 3 billion RMB for each, which is far from enough to achieve sufficient depth.

This time gap cannot be made up simply by adding more manpower. The fragmentation of the ecosystem has no visible solution in the short term.

The third bottleneck is clients.

What have domestic GPU vendors relied on to survive in the past few years? IT application innovation (Xinchuang) orders. Procurement from government agencies and state-owned enterprises is their basic market. However, the ceiling of this market is low and the budget is limited. To achieve real large-scale growth, they have to enter the commercial market.

Large internet companies, cloud service providers and large model developers do not care whether the product is domestic or not. They only care whether the card can run their business, whether the performance is sufficient, whether the ecosystem is compatible, and whether the cost performance is high.

Large model companies spend hundreds of millions of RMB on computing power every year. Choosing the wrong chip vendor means that tens of millions of RMB invested in adaptation will be wasted. Therefore, the procurement decision cannot be made only by looking at the hardware parameters.

The market is currently at the transition point between the two markets: the dividend of the IT application innovation market is fading, and the competition in the commercial market has just begun.

Most of Enflame's revenue comes from Tencent. Deep partnership is both an advantage and a risk. Once Tencent adjusts its procurement strategy, it will have a huge impact on Enflame. On the contrary, the biggest winner is system integrators such as Inspur Electronic Information, which firmly holds clients like ByteDance and Alibaba, and ranks first in the supernode market with a leading market share.

Why system integrators instead of chip companies? Because large clients need a complete set of solutions that can run their business stably. By delivering chips, network and software as a whole package, the party that is closer to the clients will have the right to speak.

According to my research, 3 billion RMB in annual revenue is a critical survival threshold.

Why 3 billion RMB? GPU companies have extremely high fixed costs. A single advanced process tape-out costs hundreds of millions of RMB, a R&D team of hundreds of people costs more than 1 billion RMB a year, and investing in the software stack costs another hundreds of millions. All these add up to 2 to 3 billion RMB of rigid annual expenditure.

If the revenue does not reach this level, it cannot even cover the daily operating costs, let alone recoup the previous R&D investment.

Moore Threads' revenue in the first half of the year was more than 1.7 billion RMB, Biren's was over 1.2 billion RMB, and MetaX's was over 1.3 billion RMB. MetaX's attributable net profit turned positive, reaching more than 600 million RMB, but it is still loss-making after deducting R&D expenses, and the other three are also in the red.

Cambricon and Hygon Information have already crossed this threshold. Their half-year revenue reached 5.9 billion RMB and 9 billion RMB respectively. Cambricon's net profit is 2.3 billion RMB. Their pricing logic is completely different from that of the four small leading vendors: they are more like infrastructure companies rather than chip startups.

The 3 billion RMB threshold not only restricts revenue, but also determines the valuation given by the capital market.

......

Moore Threads' PS ratio is about 70x, and MetaX's is 95x. In other words, the market is willing to pay 70 to 95 RMB in market value for every 1 RMB of their revenue.

Cambricon and Hygon also have premium valuations, but the logic is different. Their half-year revenue has exceeded 3 billion RMB, their business model has been proven, and their valuation is anchored on their current fundamental performance.

As for the four small leading vendors, their revenue is still far from the threshold, and their valuation is anchored on "how large you can grow in the future". To put it simply, the latter two are valued for certainty, while the four small vendors are valued for potential.

Enflame is a more representative case: its IPO price was 142 RMB, it raised 6 billion RMB, and its market value more than doubled on the first trading day. The primary market pricing is based on technical routes and order reserves, while the secondary market is directly pricing on future expectations, which shows that capital is voting with its feet.

But IPO is a one-time deal. Whether they can continue to get orders afterwards and cross the 3 billion RMB revenue threshold is the real test.

The four small leading vendors have raised a total of about 23 billion RMB through IPOs. The rapid inflow of funds shows that the capital market believes they can cross this threshold.

However, the valuations that institutions give to the four companies vary greatly.

Why is there such a big gap? In fact, the market is pricing the "certainty" of different technical routes. MetaX takes the CUDA-compatible route, which has the lowest client migration cost and the clearest commercialization path. Its bet is on "capturing the commercial market the fastest", and the 95x PS ratio is the price the market pays for this most revenue-proximate path.

Biren's chiplet + optical interconnection route has the highest ceiling, but its large-scale cluster solution is still in the verification stage, and clients are still waiting and seeing. The 38x PS ratio is the discount given by the market.

Moore Threads' full-featured GPU route has the most ambitious story, but also the highest execution difficulty. Covering all scenarios means it has to prove its capabilities from scratch, and the 70x PS ratio is a moderate level in between.

In the same sector, with the same wave of policy dividends, the valuation gap reaches 3 times, which is essentially the price difference caused by the difference in the certainty of different technical routes.

Moore Threads' market value has evaporated by more than 200 billion RMB since its peak. On the day of share lifting in early September, its stock price hit the 20% limit down, with total transaction volume exceeding 10 billion RMB in three days. MetaX's stock price hit a record low before the share lifting, falling by 40% in the past two months.

This is just the first wave. Moore Threads will have a share lifting event with 7 times the scale in December, when VC and PE shareholders will exit their positions in a concentrated manner. The combination of high valuation and large-scale share lifting will bring sustained, structural selling pressure.

The deviation between market value and financial performance has become the norm in this sector. The key question is how to distinguish between industrial options and emotional bubbles?

Just think about it, for a sector with a trillion RMB market value, what is the capital betting on? It is betting on the future of the entire track. As long as the two major logics of domestic substitution and AI computing power demand remain solid, the sector will have a solid bottom support.

In any case, GPU companies have to pass two levels of tests.

The technical level is whether they can produce qualified chips with sufficient performance; the commercial level is whether they can get enough orders to make their revenue cover all costs. They have to pass both levels at the same time, and neither can be missed.

The high valuations given by the capital market now are paying for the time window, betting that they can pass both tests within the specified time.

But the time window will not stay open forever. Moore Threads will have large-scale share lifting in December, and the four small leading vendors will enter the performance realization period one after another in 2027. The key nodes are approaching step by step, and capital's patience is not unlimited.

There will not be many companies that can finally stand out. Those with revenue exceeding 3 billion RMB will gradually follow the valuation logic of Cambricon and Hygon, and evolve into infrastructure companies. Those that cannot cross the threshold will either be acquired and integrated, or shrink back to the IT application innovation market and become regional suppliers.

No matter how high the valuation flies, it will eventually fall back to the support of revenue