More tokens, more chips: Kimi K3 embodies the "Jevons Paradox", where improved efficiency instead boosts resource consumption
Citi semiconductor analyst Peter Lee argues that Kimi K3 triggers the "Jevons Paradox": when technological efficiency gains reduce unit usage costs, total resource consumption eventually rises instead of falling, and demand for general-purpose memory such as server DDR5 and eSSD will continue to increase. Bank of America Securities believes that stronger Chinese open-source models will not necessarily prompt US AI giants to reduce their computing power investment. On the contrary, the narrowing of the model gap will raise the cost for leaders to maintain their competitive advantages.
Kimi K3 brings AI investors back to a core question: If models become cheaper and more efficient, does that mean demand for chips and memory will decline? Both Citi and Bank of America Securities lean toward a negative answer. Efficiency gains may drive far greater usage, which in turn pushes up total resource consumption.
According to Zhuifeng Trading Desk, Kimi K3 released by Moonshot AI features 2.8 trillion parameters, targeting 1 million-token context windows and long-cycle agent tasks. Citi semiconductor analyst Peter Lee believes that even if Kimi K3 is widely adopted, demand for general-purpose memory such as server DDR5 and eSSD will continue to rise.
Bank of America Securities semiconductor analyst Vivek Arya also noted in a July 17 research report that leading US AI labs are more likely to increase rather than reduce their computing power investment. If Chinese open-source models continue to catch up, leaders like OpenAI, Anthropic, and Google will need larger training scales, more intensive inference, and faster product iteration to maintain differentiation.
This means the market's pricing logic for Kimi K3 cannot focus solely on the drop in single-inference costs. More critically, the key lies in whether low-cost models drive more API calls, longer task chains, and more token generation. If the answer is yes, GPUs, HBM, DDR5, eSSD, high-speed networks, and inference systems will all likely continue to benefit.
01
Low-Cost, High-Performance Triggers the "Jevons Paradox"
Citi views Kimi K3 as a potential case of the "Jevons Paradox" in the AI industry chain. The core of this paradox is that when technological efficiency improvements reduce per-unit usage costs, usage volume may surge significantly, leading to a rise rather than a fall in total resource consumption.
The appeal of Kimi K3 comes from its combination of low cost and high performance. Public pricing shows that the cache-hit input price is $0.3 per million tokens, and the output price is $15 per million tokens. Its full weight is scheduled to be disclosed on July 27, with the goal of supporting long-cycle agent operations.
Citi's judgment is that model price reductions will not automatically reduce hardware demand. On the contrary, lower costs may increase the willingness of developers and enterprises to call models, driving more AI agent deployments. Agent tasks are not one-off question-and-answer interactions, but continuous processes that generate, read, and process tokens iteratively. The rise in call frequency and task length will translate the drop in unit cost back into higher total resource consumption.
Therefore, the impact of Kimi K3 on the semiconductor industry chain does not lie in "whether a single call is cheaper", but in "whether the total token volume expands". This is exactly why Citi is optimistic about demand for server DDR5 and eSSD.
02
Long-Context Inference Transmits Pressure to Memory
Kimi K3 adopts three key technologies: Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. Kimi Delta Attention is used to reduce the cost of 1 million-token context windows, Attention Residuals enables selective retrieval of representations across different model layers, and Stable LatentMoE enhances sparsification, activating only 16 out of 896 experts per token.
Moonshot AI's technical documentation states that this architecture enables Kimi K3 to achieve 2.5x the scaling efficiency of K2. High sparsity and long-context capabilities are the critical foundations for lowering its operating costs.
However, Citi emphasizes that this does not mean the resource pressure on the inference side disappears. Kimi K3 is still not a lightweight deployment solution: its operation requires a multi-node cluster, with super-node configurations exceeding 64 GPUs. More importantly, long-context and agent tasks will increase KV Cache usage, thereby raising the memory burden on the inference side.
KV Cache demand is directly linked to server DDR5 and eSSD. DDR5 handles high-frequency data access, while eSSD benefits from greater cache and data storage requirements. For memory manufacturers, the key variable is not whether a single model is more computationally efficient, but how many times the low-cost model is called, how many tokens are generated per call, and how long the agent task chain is extended.
03
Closing Model Gap Instead Raises the Computing Power Threshold
Bank of America Securities' conclusions echo Citi's, but focus more on GPUs and AI infrastructure. Vivek Arya argues that stronger Chinese open-source models will not necessarily prompt US AI giants to cut computing power investment. On the contrary, the narrowing of the model gap will increase the cost for leaders to maintain their competitive advantages.
Bank of America Securities points out that if open-source models continue to catch up, OpenAI, Anthropic, and Google will need to rely on larger-scale training, more reinforcement learning and synthetic data loops, more intensive test-time inference, and faster product release cadences to preserve differentiation.
This logic does not depend on which model takes the lead in the short term. Leaderboards for large models may shift rapidly, but what enterprises truly purchase is stable, low-latency, highly available AI outputs, and lower costs per unit of effective output. When model capabilities converge, competitive pressure will transmit to underlying infrastructure, including GPUs, HBM, high-speed networks, and inference systems.
Therefore, Bank of America Securities believes that the emergence of models like Kimi K3 should not be simply interpreted as "more efficient models mean fewer chips". A more realistic path is that model competition intensifies, and leaders continue to increase their computing power investment to maintain their competitive edge.
04
MoE Architecture Shifts Bottlenecks, Making VRAM and Interconnectivity More Critical
Kimi K3's adoption of the MoE architecture means that total parameter count can be decoupled from the parameters actually activated per inference. This change reduces some computational pressure, but also shifts infrastructure bottlenecks from pure computing power to memory access, expert routing, response latency, and interconnection capabilities.
Bank of America Securities emphasizes that MoE inference requires systems to move data, schedule expert modules, and connect larger computing clusters more efficiently. NVIDIA proposes that modern MoE inference requires a larger GPU domain. Its estimates show that the GB300 NVL72 can achieve up to 25x the per-watt performance of the Hopper platform on leading open-source models.
CoreWeave's tests around Kimi K2.6 also demonstrate that even if open-source MoE models activate only a subset of parameters per run, maintaining leadership in speed and cost-effectiveness still requires optimized NVIDIA GB300/GB200 NVL72 infrastructure.
This means that while model weights may gradually become commoditized, the GPUs, high-bandwidth memory, network interconnections, and inference systems required to run these models will not lose their value. On the contrary, as model call volumes expand, these segments may become areas of even more intense competition.
05
Token Usage Is Still Expanding, and Demand Has Not Peaked
Bank of America Securities also cites OpenRouter data indicating that model call volumes are still growing rapidly. As a third-party API platform connecting multiple model labs, token usage on OpenRouter continues to rise, with token consumption from Chinese AI lab models already exceeding that of non-Chinese lab models.
This data does not directly represent all enterprise and consumer scenarios, but it at least shows that developers' choices in the open model ecosystem are changing. Falling model prices and improving capabilities will likely drive further expansion in call volumes, rather than being offset by the efficiency gains of a single model.
Enterprise paid usage is also on the rise. Ramp data shows that as of June 2026, approximately 55% of US enterprises have paid subscriptions for AI models, platforms, or tools, up from the 21% estimate in the US Census Bureau's BTOS survey. By model, Anthropic has an enterprise adoption rate of 42.4%, while OpenAI's rate stands at 39.5%.
However, AI spending remains highly concentrated. The top 1% of enterprise users spend an average of about $4,833 per employee per month on AI, the top 10% spend $516, while the overall median is only $11. The technology and media industry has a subscription rate of 79.8%, and large enterprises have an adoption rate of 65.5%, higher than 61.3% for medium-sized enterprises and 48.7% for small enterprises.
This set of data supports the judgment that AI usage is still in the diffusion phase. If low-cost models lower the barriers to entry, future incremental demand will likely come from more enterprises, more developers, and more agent applications.
06
Shifting from "Saving Computing Power" to "More Workloads"
The shared conclusion of Citi and Bank of America Securities is that the significance of Kimi K3 lies not in whether a single model reduces per-unit inference costs, but in whether low costs unlock larger-scale workloads.
For Citi, the most direct beneficiary is server DDR5 and eSSD. Long-context, KV Cache, and agent tasks will increase memory access and storage demand. For Bank of America Securities, the more critical segments are GPUs, HBM, high-speed networks, and inference systems, as model competition will force leaders to continue investing in infrastructure.
Bank of America Securities also mentioned that the "45nm open-source EDA" related to Kimi K3 should not be simply interpreted as a replacement for commercial EDA. On the contrary, it indicates that chip design still relies heavily on EDA tools. For advanced process nodes, commercial EDA vendors such as Cadence and Synopsys remain in critical positions.
This article is from the WeChat public account "Hard AI", author: ZHAO Ying, editor: Hard AI, published with authorization from 36Kr.