Hardkr Exclusive | Heterogeneous Intelligent Computing Platform Raises Tens of Millions Yuan in Angel Round Financing, Targeting NVIDIA and Domestic GPU Computing Power Scheduling
Hard Krypton learned recently that Bolun Zhihui, a heterogeneous intelligent computing platform and gradient Token factory technology service provider, has completed a multi-million-yuan angel round of financing, with a post-investment valuation of about 150 million yuan. This round of financing is invested by a first-tier industrial fund that focuses on the AI and AI Infra tracks and has previously deployed large model training frameworks. The funds will be mainly used for the iteration of core heterogeneous intelligent computing technologies, the polishing of computing power infrastructure products, and the expansion of core technical teams.
Bolun Zhihui was founded in 2026 and is headquartered in Beijing. Liu Jinzhi, the CEO, once worked at Intel Asia-Pacific R&D Center and served as the China Head of Silicon Valley Distributed System Software, with rich experience in the fields of computing architecture and system software. In 2016, Liu Jinzhi was a core member of the world's first batch of AI startups focusing on natural language processing, and later participated in the construction of China's first batch of Nvidia Superpod training clusters to support the R&D and service implementation of large language models.
Yang Guang, the CTO, has successively been in charge of software R&D and system architecture in multinational enterprises such as Siemens Communications, Oracle, IBM and Huawei Cloud, and is an early member of cloud-native R&D. All core members of the company come from innovative enterprises including IBM, Iluvatar CoreX, Huawei and DriveStack. The core team has previously operated the first batch of domestic Nvidia SuperPod training clusters, and realized the inference of 10,000 domestic heterogeneous GPUs at a relatively early stage.
With the construction boom of intelligent computing centers across the country, the nominal computing power scale continues to rise, but many cabinets are in an "idling" state after being put on shelves. Having chips does not mean having computing power, and having computing power does not mean it is available. From chips to output Tokens, there are several engineering gaps in the middle. How to uniformly manage chips of different architectures? How to schedule and optimize fragmented computing power? How to grade computing power on demand to produce different values?
At the same time, the global daily call volume of AI Tokens has exceeded 300 trillion times. After the rise of Agents, the complex tasks of intelligent bodies have increased, and the amount of Tokens consumed by a single task often reaches 100,000 or even millions of levels, leading to a sharp increase in computing power demand. The global AI inference chip market size has reached nearly 50 billion US dollars, and inference costs have jumped to the largest operating expense for AI enterprises.
Liu Jinzhi introduced to Hard Krypton that based on this background, the company has independently developed the TOLD (Token Oriented Large-scale Distributed Architecture) technology stack, which optimizes computing power into stepped Token output.
Liu Jinzhi entered the intelligent computing track back in 2023. "Before the ban on high-end Nvidia computing power, we built the first batch of domestic Nvidia SuperPod, that is, the SuperPod architecture based on Nvidia DGX, and carried out in-depth cooperation with leading large model companies." On the other hand, he led the team to get involved in the ecosystem and technical system of domestic GPUs at a relatively early time.
Affected by various factors such as geopolitics and the rise of domestic GPUs, the GPU assets held by various enterprises are very scattered, including both Nvidia GPUs and different types of domestic GPUs. "The problem is that these different GPUs cannot be directly used in coordination due to different technical routes and other reasons. It is a technical difficulty to both coordinate them and achieve optimal performance."
Therefore, how domestic enterprises integrate and optimize these heterogeneous computing power assets has become a rigid demand. "Moreover, as the GPU assets accumulated by various enterprises, state-owned enterprises and central enterprises become more and more abundant and diverse, this will become a very large demand."
Bolun Zhihui's self-developed architecture TOLD is compatible with Nvidia and multiple domestic GPU architectures. Its self-developed cloud-native distributed scheduling technology realizes unified management and inference scheduling of 10,000-card level heterogeneous GPUs, enabling different chips to work collaboratively with optimal model efficiency under the same framework. For tokens and applications, it can also provide one-click adaptation for large models and the B2B2C multi-tenant innovation mode.
Liu Jinzhi introduced that the company already has signed customers at present, and will launch hardware products later. The following is an excerpt of the interview between Hard Krypton and him:
Hard Krypton: In the current large model stage, what are the main links that computing power is used for?
Liu Jinzhi: From the perspective of the large model stage, computing power is mainly divided into two parts: large model training, and inference applications after the training is completed.
Now large model companies are relatively concentrated, and some state-owned and central enterprises are also training large models on their own, but they will not commercialize them externally, only use them internally, and their investment is also very large. These companies are still continuously laying a large number of training computing power. At present, China's high-quality training computing power is less than one-tenth of that of the United States, so the training computing power still needs to catch up.
The larger market in the future is inference. DeepSeek R1 has driven the real implementation of inference and AI. Coupled with some new applications emerging this year, it has driven a large number of inference demands for Token implementation. In the long run, the demand for inference may reach more than 10 times or even 100 times of the training demand. With the implementation of a large number of AI applications, inference will theoretically become a more extensive market, and it is also a place where domestic computing power can play a great role.
Hard Krypton: What is your own TOLD technology stack specifically? Where is the technical barrier reflected?
Liu Jinzhi: It has several characteristics. The first is Token-oriented, to make Token achieve the fastest and optimal performance. The second is Large Scale. Many companies in the market are talking about heterogeneity, but our feature is that we have actually built large-scale domestic GPU clusters. From an engineering point of view, large-scale heterogeneity is completely different from deploying several heterogeneous cards. For example, the performance of some latest domestic GPUs has been close to Nvidia H200, but large model training requires thousands or even tens of thousands of cards to form a super-large-scale cluster, and the stability of such a cluster is the biggest challenge.
In addition, our clusters can be distributed across the country or even around the world. In fact, we have managed more than 10 computing centers at the 1000-card level distributed all over the country, reaching the 10,000-card level. In this process, we have done a lot of R&D work. The first layer is the scheduling layer. We are based on cloud-native, but we will deeply modify some underlying contents of cloud-native. It is not a single architecture. We need to integrate cloud, distributed architecture, large-scale distributed architecture, and the understanding of the model itself into the product.
At present, it has supported more than 5 types of heterogeneous GPUs, including large-scale GPUs of Nvidia and domestic brands.
Hard Krypton: Has this large-scale heterogeneous platform been verified at present? What scenarios is it mainly used in?
Liu Jinzhi: We have actually built a 10,000-card level domestic GPU cluster composed of more than 10 1000-card level clusters, similar to a "large Internet cafe", and this cluster provides services to the public in real scenarios. Later, some cards were rented out commercially, but we have completed the verification on it.
This scenario is actually quite versatile, which can be for B-end and also for C-end. For the B-end, it can provide internal services for enterprises. For example, a company that makes AIGC short dramas needs to call different language models, multimodal models and image generation models, and it can also be used for embodied intelligence or other pan-ecological scenarios. For the C-end, it is for personal use. For example, personal "lobster" (Agent) raising, Agent tool calls, codex, claude code and other tools can all access the models we provide.
Hard Krypton: Why do you value the C-end? What kind of business model will you adopt later?
Liu Jinzhi: Now everyone is saying that they cannot afford to use "lobster", because it needs to run tasks in standby day and night. The core problem is that there is no good stratification at present. Everyone wants to use the top model together, and the top model also means that the most expensive cards are used behind it.
But in fact, a lot of work, such as sorting out files and data, can be completely done with basic models, which means that cheaper cards are sufficient, especially domestic GPUs.
In terms of business model, for the C-end, we will first operate part of the Token factory by ourselves, which is built on our own platform. We will do a good job of debugging to ensure that users can choose the model they want and get the optimal solution at the same time. For the B-end, we can cooperate with enterprises to provide services for their own private intelligent computing clusters, and even design from scratch to see what kind of card type matches to meet their specific scenarios.