Don't even think about deploying open-source large models without two Rolls-Royce Phantoms.
Ever since Kimi K3 went open-source, "deploying large language models" has suddenly become a hotly debated topic across the internet.
Ranked the third in global performance, available for free download. Who could resist such an offer?
But according to the most widely circulated claim online, getting K3 up and running on your own would cost roughly 30 million RMB on equipment. After all, K3 has a total parameter count of 2.8 trillion, its weight files alone take up 1.4TB to 1.5TB of storage, and the actual video memory required exceeds 2TB.
Therefore, the official recommendation from the Kimi team is that at least one supernode with 64 accelerator cards is required to self-host Kimi K3. Calculated based on the current mainstream B200 model, one card costs around 400,000 RMB, so 64 graphics cards alone add up to 25.6 million RMB. The floor of an ordinary residential building cannot support equipment that heavy, so you also need to consider building a dedicated computer room, plus the electricity costs for equipment operation and supporting liquid cooling systems. To put it simply, if you are dead set on running Kimi K3 on your own, the upfront investment would be enough to buy no less than two Rolls-Royce Phantoms.
However, even if you have extraordinary connections and manage to get all the equipment sorted out, deploying Kimi K3 personally is far from easy, because you will still face subsequent costs including electricity bills, heat dissipation, model tuning, operation and maintenance, all of which are considerable expenses.
Besides K3, there are many other domestic open-source large language models, such as DeepSeek V4-Pro, GLM-5.2, and Qwen 3.8, all of which are players with hundreds of billions or even trillions of parameters. So what is their deployment cost?
How Large Are Open-Source Large Language Models Exactly?
In fact, the idea that "ordinary people want to run large language models on their own" did not just emerge recently. It is only because K3 became so viral that the topic was brought up again by major media platforms.
Back when DeepSeek V3 was released, many people were already discussing how ordinary users could deploy large language models.
At the end of 2024, DeepSeek V3 was open-sourced with a MoE architecture of 671 billion total parameters and 37 billion activated parameters. Its training cost was only 5.57 million USD, while its performance was on par with Claude 3.5 Sonnet.
For a while, Zhihu, Bilibili, and WeChat Official Accounts were flooded with tutorials that "walk you through local deployment of DeepSeek V3 step by step".
Some people pushed their RTX 4090 to the limit, some tripped their home electricity meters, and even relatives with no technical background were asking "What is DeepSeek and how do I use it".
The scene was extremely similar to the period when the Bitcoin market took off, with everyone talking about mining with their laptops.
But for all the hype, to run V3 fully, the recommended configuration is 8 H100 or A100 cards. The official price of a single H100 card is 36,500 USD, equivalent to about 260,000 RMB. The A100 is cheaper, with a single card costing around 108,000 RMB, so an 8-card server costs no less than 800,000 RMB. It is worth noting that the top-spec BMW X5, Mercedes-Benz GLE, and Audi Q7 only cost around 700,000 RMB each after tax and licensing.
In addition, those so-called "local deployment with RTX 4090" tutorials online only run a 4-bit precision stripped-down version that has been quantized beyond recognition. The output answers are not only low in quality, but also have strict word count limits. Once the content exceeds a few hundred words, the model will crash instantly.
The gap between ordinary users and trillion-parameter models is far larger than the "5-minute setup" described in those tutorials.
By 2026, the situation has become even more exaggerated.
DeepSeek V4-Pro, released on April 24, 2026, has a total parameter count of 1.6 trillion, is open-sourced under the MIT license, and its weights are available for free download on Hugging Face.
Even if quantized to FP8 precision, the weight file starts at 1.6TB. To fit it fully into video memory, at least 16 H200 cards are required. A single H200 card has 141GB of video memory, so 16 cards add up to 2.2TB, which is just enough after deducting KV cache and system overhead.
The official NVIDIA price for a single H200 card is 27,000 USD, equivalent to about 190,000 RMB, but the domestic quotation is 235,000 RMB. An 8-card server is quoted at 1.88 million RMB, so 16 cards cost 3.76 million RMB.
Then there is GLM-5.2, open-sourced by Zhipu AI on June 17, 2026, with a total parameter count of about 753 billion and a smaller number of activated parameters. It also only requires 8 H200 cards to run.
Finally, there is Qwen 3.8. Alibaba released its preview version on July 19, 2026, with a total parameter count of 2.4 trillion, and promised to open its weights soon.
To fully accommodate 2.4 trillion parameters, the 141GB video memory of the H200 is no longer sufficient. You need to use the B200, which has 192GB of video memory per card, and 64 B200 cards add up to 12.3TB to support this scale. In other words, to deploy Qwen 3.8, you also need to prepare 30 million RMB, just like the K3 mentioned earlier.
However, this is not the end of the story, it is only just beginning. The graphics cards are just the first expense. The next cost is electricity.
Take the 64 B200 cards for Qwen 3.8 as an example. The TDP of a single B200 card is 1000W, and the liquid-cooled version can reach 1200W. 64 cards under full load consume 64kW to 77kW of power. Adding the power consumption of surrounding components such as CPUs, network switches, and cooling pumps, the full rack easily exceeds 80kW.
The heating power of a 1.5-horsepower household air conditioner is about 1200W, so a single B200 card almost consumes as much power as an air conditioner. 64 B200 cards running at full load are equivalent to turning on more than 50 air conditioners at the same time, which is equal to the power consumption of half a residential building.
Calculated at 0.8 RMB per kWh for industrial electricity, 80kW running at full load costs 1536 RMB per day, which is 560,000 RMB per year. This is the case of non-stop operation. Once a large language model inference server goes online, it basically runs 24/7, with no option to "turn it off to save power".
After the power issue is addressed, the heat issue comes right next.
The 1000W power consumption of a single B200 card means that traditional air cooling cannot handle the heat dissipation at all.
NVIDIA officially stated that the SXM version of the B200 was designed for liquid cooling from the very beginning. The performance upper limit of the air-cooled version is only 18 PFLOPS, while the liquid-cooled version can reach 20 PFLOPS.
Liquid cooling equipment can be divided into two types: the first is NVIDIA-certified, with a market price of about 150,000 to 200,000 RMB per cabinet, and the second is non-NVIDIA-certified, priced at about 80,000 to 150,000 RMB per cabinet.
Space is also a problem. A standard 8-card HGX server is 6U to 8U in height, where U is the standard rack unit, and 1U equals 44.45 mm, so the height can be understood as 0.267 meters to 0.356 meters.
The bare server generally weighs 75–120 kg, and when fully configured with rails, cables, and PDUs, the whole device usually weighs more than 90 kg. But the base area of a single cabinet is very small, usually 0.36 square meters. The uniform live load of ordinary residential buildings is 200 kg per square meter, which means that the local pressure of a cabinet is far higher than the conventional civil design standard, so it cannot be placed in an ordinary residence.
In addition, this kind of equipment has strict requirements for ambient temperature, humidity, and dust. If these conditions are not properly controlled, there will be risks including coolant leakage.
Finally, there is operation and maintenance. Large language model deployment is not as simple as "installing a software".
Model weight loading, inference framework tuning, KV cache policy configuration, monitoring and alerting, log troubleshooting, and fault recovery all require professional engineers. At least a small team of 3 to 5 people is needed. According to the current market, a qualified AI infrastructure engineer has an annual salary of no less than 300,000 RMB, so the annual labor cost of such a small team will exceed one million RMB.
Money Cannot Solve All Problems
Price is only one aspect. The following points are what really make people feel powerless.
It is no longer news that NVIDIA's high-end cards are subject to export controls to China.
The U.S. Bureau of Industry and Security (BIS) explicitly stipulates that B200, B300 based on the Blackwell architecture, and the next-generation Rubin series chips will still be strictly prohibited from being exported to China for at least 18 to 24 months after their global release.
Fortunately, since we only use the model for inference instead of training large language models, we only need to consider the inference performance of the GPU. Domestic AI chips can also handle inference tasks.
But then again, being able to run is one thing, running well is another. NVIDIA's CUDA ecosystem has been accumulated for more than ten years, with millions of developers worldwide, and there are countless answers you can find on Stack Overflow.
But for domestic AI chips, although DeepSeek V4 has achieved full-stack domestic training, 8 major domestic chip manufacturers have completed Day 0 adaptation, and Kimi also announced that it will support running on domestic cards in the future, these are the engineering capabilities of the official DeepSeek and Kimi teams.
As a civilian player who gets a domestic card, the task that can be completed on CUDA in one day will take you a long time to get adapted on a domestic card. When an error pops up, you cannot find a solution on Stack Overflow.
The gap in the software ecosystem is more frustrating than the gap in hardware performance. If the hardware is slow, you can wait. If the software cannot run, you do not even have the qualification to wait.
To say the least, even if you have special channels to get the cards, set up the liquid cooling system, afford the electricity bills, maintain the team, and get the model running — what then?
When you compare this self-built system with the official API, you will find that 99% of companies can never recoup their investment from self-hosting.
Let's do the math with the minimum configuration scheme for DeepSeek V4-Pro.
The API pricing for DeepSeek V4-Pro is 3 RMB per million input tokens (only 0.025 RMB when cache hits) and 6 RMB per million output tokens. It has also pioneered "peak-valley time-of-use billing", which cuts the computing cost by 60% during off-peak hours.
Without considering the difference between peak and valley prices, assume your team calls 50 million tokens per day, including 35 million input tokens and 15 million output tokens, which is already a very high usage volume.
The input cost is 35 * 3 = 105 RMB per day; the output cost is 15 * 6 = 90 RMB per day. The total cost is less than 200 RMB per day, which is about 70,000 RMB per year.
As mentioned earlier, 16 H200 cards alone cost 3.76 million RMB, plus 300,000 RMB for liquid cooling transformation, 300,000 RMB for power capacity expansion and site renovation, the first-year hardware investment is about 4.36 million RMB.
Calculating the operation and maintenance cost, a small team costs one million RMB per year in labor. For 16 H200 cards running at full load with a power consumption of about 11.2kW for 24 hours a day, plus the power consumption of the CPU, memory, hard disk, and fans of the whole server, and the PUE of the data center, the actual power consumption is about 15-20kW. Calculated at 1 RMB per kWh for industrial electricity, the annual electricity cost is between 130,000 RMB and 180,000 RMB.
In other words, without counting hardware depreciation, just "keeping this system running" costs more than 1.2 million RMB per year. If the hardware is depreciated over 5 years, the total annual cost of self-hosting exceeds 2 million RMB.
The core logic of large language model inference cost is that the larger the model size, the more expensive the inference is. However, API vendors' ability to dilute costs through economies of scale far exceeds that of any single enterprise.
Vendors such as DeepSeek, Kimi, and Alibaba run hundreds of user requests on a single GPU. Batch processing, KV cache reuse, speculative decoding, and dynamic routing, every technology is used to squeeze out every bit of computing power.
Kimi Kimi K3 has a cache hit rate of over 90%, and the input cost after a hit is only a quarter of the standard price. If you deploy it on your own, one card only serves you, and the GPU utilization may not even reach 10%.
Using the API honestly is not a shame.
The Model You Downloaded Is Not the Same as the Official Model
Even if you have special channels to get all the required equipment, and you spend several all-nighters finally successfully deploying the desired large language model on your own device.
Then congratulations, the thing you are running is not the same species as the one in the official API at all.
Why? Because what you can download is only the weight file.
The weight file is the "memory" of the model. Everything the model learned during the training process, such as language rules, knowledge, logic, and even a certain degree of "common sense", exists in these weights in the form of numbers.
However, the model weight is not a database, it is the way the large language model understands problems. You can think of model weights as a kind of "importance score".
How to determine if there is a cat in a photo? You might look at the face, whether there are whiskers, whether there is fur on the body, whether there is a tail...
The face is the most important. Felines have similar faces, but whether there is a tail