HomeArticle

After the 2.8 trillion-parameter era, why is Silicon Valley re-aligning with the open source camp?

36氪的朋友们2026-07-28 11:44
The "Kubernetes moment" of open models

On July 24, Jensen Huang posted his first X post.

He shared an open letter titled "Open Weights and U.S. AI Leadership." The signatory list initially included only 25 companies and organizations such as NVIDIA, Microsoft, Meta, and IBM, but quickly expanded to around 50, with OpenAI, Google, AMD, Cisco, Cloudflare, and GitHub joining later. Among major leading U.S. model companies, Anthropic became the only one that had not yet signed.

Jensen Huang's caption was just one sentence: "The world needs both leading closed-source models and leading open models."

A few days earlier, he expressed a more direct stance in an interview with foreign media. Jensen Huang stated that China's open-source models are "very excellent," and open-source models with strong performance should be used.

This collective statement took place after a key milestone. In mid-July, Moonshot AI released Kimi K3. According to official disclosures, this model adopts a Mixture-of-Experts architecture with a total of 2.8 trillion parameters, 896 expert modules, activating only 16 of them per Token, and supports a 1 million-token context window.

Kimi K3 thus became one of the largest publicly available open-weight models by parameter scale. What has drawn more attention is that it has entered the global first tier in multiple evaluations covering agents, programming, and front-end development.

A Chinese system that can be downloaded and deployed, with performance close to top-tier closed-source models, has quickly come into the view of global developers. Along with it, renewed discussions have emerged about the model moats of U.S. AI companies, the competitiveness of open ecosystems, and how Washington should treat Chinese models.

On one side are government concerns about technology diffusion, model distillation, and national security; on the other side are the practical needs of tech companies for model supply, cost control, and innovation space. Jensen Huang's statement also reflects the core of this game.

01 Open Models Enter the Frontier Capability Zone

Open-weight models have long been regarded as a supplementary option to closed-source flagship models: lower in price, more flexible in deployment, but often with a noticeable gap in overall capabilities.

Kimi K3 has changed this perception.

The 2.8 trillion parameters do not mean that such a huge amount of computing power is called for in every inference. The Mixture-of-Experts architecture selects a small number of experts to participate in computation based on the task, expanding the model's knowledge capacity while controlling the cost of a single inference.

Kimi K3 also adopts technologies such as mixed linear attention, attention residuals, and low-precision quantization to reduce the resource pressure brought by long contexts and large-scale inference.

On July 27, Moonshot AI officially released K3's full technical report and model weights. The report reveals the specific implementation of the above technologies: KDA (Kimi Delta Attention) mixes linear attention and full attention layers at a 3:1 ratio, with the former responsible for efficiently processing long sequences and the latter maintaining precision at key positions, similar to the division of labor between key frames and differential frames in video coding.

Attention Residuals establish direct connections between network layers to compensate for the information loss caused by linear attention; Stable LatentMoE uses statistical methods to balance the training load of 896 experts, enabling the architecture with extreme sparsity (activating only 16 experts at a time) to converge stably.

The report also discloses three pieces of training infrastructure: the expert communication library MoonEP, the high-performance operator FlashKDA, and AgentEnv, an agent sandbox that supports million-token-level reinforcement learning. Moonshot AI claims that this architecture improves the model's intelligence level by about 2.5 times under the same computing power. This is also why K3, despite having 2.8 trillion total parameters, keeps its actual inference cost within an acceptable range for enterprises.

What makes Kimi K3 truly industry-relevant is that large-scale open models have begun to meet three conditions at the same time: performance close to leading closed-source models, downloadable and modifiable weights, and inference costs that can be controlled within an enterprise's affordable range.

This gives enterprises new options.

General office, content generation, and coding tasks can continue to use closed-source APIs; when it comes to internal data, core code, cybersecurity, and compliance requirements, enterprises can deploy open models on their own servers or cloud environments, independently controlling data flow, access permissions, and model versions.

According to reports, Microsoft is evaluating the use of Kimi K3 for Azure and some Copilot scenarios to reduce its dependence on high-cost closed-source models. The externally circulated claim that "up to $600 million can be saved annually" has not been publicly confirmed by Microsoft, but it reflects a trend: open models are becoming tools for large tech companies to control inference costs and enhance their bargaining power with suppliers.

Open-weight models have thus evolved from a low-cost "affordable alternative" to a frontier capability zone that can influence enterprise technical architectures and procurement decisions.

This is also the important background for Jensen Huang to choose to speak out at this time.

02 Silicon Valley Realigns Around Open-Source Models

After Kimi K3's release, debates over open-source models in the U.S. quickly intensified.

The pro-restriction side worries that the widespread distribution of advanced model weights could be used in high-risk applications such as cyberattacks and biological research. The U.S. government is also concerned about whether Chinese models acquire U.S. model capabilities through distillation, and whether related companies use advanced chips subject to export controls.

Both OpenAI and Anthropic previously warned Washington that Chinese open-source models are threatening U.S. AI leadership, and called on the government to strengthen regulation over model distillation, chip usage, and technology diffusion.

The pro-open-source side, however, argues that model sources and deployment methods need to be discussed separately. By downloading model weights to local devices for operation, enterprises can isolate external connections, control access permissions, and retain data. For highly sensitive scenarios such as cybersecurity, finance, healthcare, and government, the ability to deploy, audit, and modify models privately is itself part of security capabilities.

Nearly 200 U.S. startups previously wrote to the government to oppose a full ban on Chinese open-source models. They are concerned that excessive restrictions will raise model costs for startups and further strengthen the monopoly of a small number of U.S. closed-source model suppliers.

A new interest structure has thus taken shape in Silicon Valley.

NVIDIA sells GPUs and a full range of AI infrastructure; the richer the open-source models, the stronger the enterprises' demand for model deployment. Microsoft needs to operate cloud computing, model APIs, and application products at the same time, and introducing more models helps reduce costs and enhance bargaining power. Meta has long relied on the open ecosystem to expand its influence among global developers. Companies such as IBM, Cisco, and Cloudflare focus more on the private deployment, hybrid cloud, and cybersecurity markets.

Open-source models have expanded the number of participants in the entire AI market, as well as the potential customer base for infrastructure suppliers.

Anthropic's stance is noticeably different. In the joint open letter, Anthropic has continuously been absent. It has long taken leading model safety as its core strategy, and warned that open advanced model weights could lower the threshold for the diffusion of risks.

This choice is highly consistent with Anthropic's brand positioning and closed-source business model. The stronger open-weight models become, the lower enterprises' dependence on high-cost closed-source APIs will be, and the greater the pressure on the safety narrative and model services that Anthropic relies on to build differentiated competition.

A technical route debate over open source vs. closed source is, in fact, backed by more complex negotiations over industrial interests and policy boundaries.

03 What Jensen Huang Is Really Defending Is the Computing Power Market

Jensen Huang's support for open-source models follows a clear business logic.

The richer the model supply and the lower the usage cost, the more people and enterprises will enter the AI market, and the greater the total number of Tokens generated and the demand for computing power will eventually be.

After DeepSeek's release in 2025, the market once worried that efficient models would weaken GPU demand. Jensen Huang has responded multiple times that improved model efficiency will expand the scope of usage: as single inference becomes cheaper, developers will call more models, run them for longer periods, and introduce AI into more tasks.

This aligns with the Jevons Paradox. When the efficiency of using a resource increases and its unit cost decreases, total demand may grow accordingly.

AI computing demand has also gone beyond mere pre-training. Models need to go through supervised fine-tuning, reinforcement learning, long-running inference, and agent execution. For an agent to complete a complex task, it may need to continuously call models, search for information, operate tools, verify results, and re-plan based on feedback.

Every step consumes Tokens.

Jensen Huang therefore describes data centers as "AI factories": input electricity and data, output Tokens. GPUs, networks, storage, power, and data centers form an industrialized computing system.

In this logic, the business models of model companies can keep changing, while computing power remains the common foundation.

Closed-source models generate revenue through APIs, subscriptions, and enterprise services; open-source models form a business ecosystem through cloud deployment, fine-tuning, hosting, and industry-specific solutions. Both paths require training clusters, inference servers, high-speed interconnection, and software toolchains.

Open-source models will also bring more SMEs into this market. Enterprises may start experimenting with an open-source model, and then purchase cloud computing power, model hosting, security services, and more stable commercial solutions. As models enter production environments, reliability, latency, concurrency, permissions, and operation and maintenance will all generate new infrastructure demands.

What Jensen Huang values is exactly this ever-expanding market base.

04 After 2.8 Trillion Parameters, Competition Shifts to System Capabilities

Kimi K3 has pushed open models to 2.8 trillion parameters, but the Mixture-of-Experts architecture gradually decouples total parameters from actual computation. The number of parameters a model has only reflects part of the system's capacity. Activated parameters, training data, post-training methods, inference efficiency, tool usage capabilities, and deployment costs are collectively determining the commercial value of models.

The next stage of competition will focus on three levels.

First is efficiency. The core indicators that enterprises care about are how many Tokens, how many rounds of calls, and how much computing power are needed to complete a task. A model scoring high in evaluations does not necessarily mean that its actual usage cost is lower. Capability, speed, and cost need to be measured simultaneously.

Second is the software ecosystem. Around open models, developers are building quantization tools, inference frameworks, chip adaptation, professional fine-tuning, model fusion, and agent systems. Hugging Face has hosted over 2 million public models, and competition between individual models is evolving into competition between model families and ecosystems.

Finally, there is deployment control. More and more countries and enterprises hope to have their own models, data, and computing power infrastructure. Open-weight models can be localized, modified, and run on different chips, making them an important technical foundation for sovereign AI and enterprise private deployment.

Some call this phase the "Kubernetes Moment" for open models.

Kubernetes was originally open-sourced by Google, and then gradually became an important standard for cloud computing infrastructure with the joint efforts of developers, cloud vendors, and enterprises. Open models may also evolve along a similar path: model weights become the base layer, surrounded by growing inference frameworks, cloud services, hardware adaptation, and industry-specific applications.

It remains highly uncertain whether the open ecosystem can form a unified standard similar to Kubernetes. Model architectures, licenses, inference frameworks, and hardware platforms are still highly fragmented.

But the direction is gradually becoming clear: when open models enter the frontier capability zone, closed-source companies' competitive moats will shift more to product experience, data closed loops, service stability, and professional scenarios; open models will expand the market by leveraging cost advantages, control, and ecosystem innovation.

Jensen Huang's statement is not a show of support for "Chinese open-source models." No matter whether the U.S. closed-source models, Chinese open-source models, or enterprise-specific fine-tuned industry models eventually win, as long as AI applications continue to expand and more systems begin to generate Tokens continuously, NVIDIA will remain in the most stable segment of industrial growth.

From this perspective, Jensen Huang's first X post is also a business proposal letter to Washington. What he is defending is an AI market with more diverse model supply, a larger number of developers, and continuously expanding computing power demand.

This article is from the WeChat public account "Tencent Tech", written by Tencent Tech, and published by 36Kr with authorization.