The era of making easy passive profits by relying on APIs is coming to an end: Kimi has smashed the traditional industry moats, OpenAI is left completely isolated with no external support, and Google is still deeply mired in internal friction.
Using company-issued computers to conspire to steal their former employer's roadmap in office software, even typing out "LOL" — this likely most technically incompetent leak case in history has just been caught red-handed by Apple. Dozens of OpenAI employees have received legal preservation notices.
Behind this absurd lawsuit lies OpenAI's increasingly isolated situation. Elon Musk, Anthropic founder Dario, Microsoft — their biggest financial backer who has been subtly mocking them for misusing user data — and now Apple as well. Former friends are turning into enemies one by one.
Not only are they losing friends, their core profit-making logic is also on the verge of being overturned. Previously, major tech players followed a straightforward playbook: pour massive funds to train the most capable models, then charge exorbitant API fees. But now this moat is eroding. Moonshot just launched the 2.8-trillion-parameter Kimi K3 last week, whose overall performance has steadily landed in the range between GPT-5.5 and Opus 4.8. The generation gap between open-source models and the top-tier US models has been drastically shortened from 15 months to just 4 months. Facing reliable alternatives that are 40% or even 70% cheaper, enterprises can hardly keep paying high premiums for "cutting-edge intelligence" with their eyes closed.
Microsoft's Satya recently mentioned the "Reverse Information Paradox" online, warning enterprises not to hand over all their core data just to use AI models, which would only make service providers stronger and stronger. This is essentially a subtle hint to their customers: don't buy APIs from those firms.
In the latest episode of the *Big Technology Podcast*, guest Ranjan Roy and the host broke down this chaotic week:
The model of charging premium API fees by monopolizing top-tier intelligence is no longer viable. The Kimi K3 proves that open-source models have shortened the technical gap to four or five months, with performance firmly in the first tier. When large enterprises and cloud providers can directly deploy open-weight models at this level, the high profits of cutting-edge models will be drastically squeezed, and intelligence is starting to become a commodity.
The competition focus has shifted to "who can build better products". Since everyone has access to highly capable models and affordable alternatives, cutting-edge labs can no longer expect to profit from model sales forever. If you bet on OpenAI and Anthropic to win, you have to believe they are the world's best AI product teams — and Anthropic is currently proving this with Claude Code.
Google is even renting out its computing power, but its flagship model has been delayed due to internal turf wars. Google still has advantages in multimodal capabilities, but moves very slowly in AI programming. Google Cloud, DeepMind, and Android are all developing their own programming tools, and even some purist engineers are resisting using AI to write core code — internal friction has drastically dragged down progress.
Behind that "dumbest leak case" is the fact that OpenAI has no friends left. OpenAI has poached roughly 400 former Apple employees, and now dozens of them are caught up in the lawsuit. Not just Apple, Microsoft's Satya is also using the "Reverse Information Paradox" as a defensive move. In this industry, you can't go far by working alone.
The Eve of the Falling Walls: How Did Kimi K3 Shock Silicon Valley and Wall Street at the Same Time?
Host: The walls around OpenAI and Anthropic are collapsing: cheaper models, plus new Chinese open-source competitors, are challenging their profitability. Meanwhile, Google has delayed its flagship model, seemingly spinning its wheels in place. Also, why can't OpenAI retain its partners, but instead make them feel dissatisfied one after another?
This episode belongs to Kimi K3. We will fully discuss this new challenger from China: a 2.8-trillion-parameter model that is pushing the boundaries of cutting-edge AI, and has outperformed almost all models except Fable. Then we will talk about what this means, and whether the price war has truly broken out, especially after Meta and SpaceX also released cheaper models.
We will also discuss the current situation at Google: what exactly went wrong there, and why its latest model was delayed. And of course, OpenAI — we didn't even get a chance to cover it last week, because just hours after last week's podcast was released, Apple announced it was suing OpenAI. So today we will definitely talk about this lawsuit, and more importantly, why OpenAI can no longer retain its partners (at least not for long).
As always every week, we are joined by Ranjan Roy from Margins. Ranjan, great to see you.
Ranjan Roy: Happy Kimi K3 Day. This is the Kimi Moment. Ready?
Host: Absolutely ready. This is an extremely significant week for the AI story, and possibly a game-changing week, because Kimi K3 has joined the conversation.
Let me start by reading a Bloomberg report, then we can dive in. Bloomberg writes: "Powerful new Chinese AI surprises investors, spurs tech stock sell-off. An unexpected breakthrough from Chinese AI startup Moonshot rippled through global markets on Friday. Investors are comparing it to last year's DeepSeek moment, triggering sharp drops in AI and semiconductor stocks. The trigger was Moonshot's new model Kimi K3, which the company says can rival the strongest products from OpenAI and Anthropic. The release was quickly dubbed the new Kimi Moment."
Venture capital analyst Ling, who is managing director at Union Square Ventures, made an interesting point. She says people are worried that if US companies start using more Chinese models and fewer Anthropic models, Anthropic will reduce its investments. This means US companies will cut capital expenditures, and eventually chip demand will be affected too.
The situation is roughly this: OpenAI and Anthropic's models have been leading the pack, while a batch of other models lagged roughly 10 to 15 months behind, suitable for handling less demanding tasks. But Moonshot, which previously launched the highly successful Kimi K2, has now released this massive 2.8-trillion-parameter model Kimi K3. It has matched GPT-5.6 and Opus 4.8 on almost all benchmarks. It even beat Fable 5 on ProgramBench and SWE Marathon. By all the benchmarks we use to measure model capabilities, it has secured a spot among the top-tier models.
This matters because it shows Chinese open-source models may not be 10 to 15 months behind US AI, but only 4 or 5 months at most, or even less. If that's the case, why choose closed, expensive models over open-source ones? When others can catch up this quickly, how much value is there in being the first to build the strongest model? That's my take. Ranjan, what do you think?
Ranjan Roy: I think this is a huge event, truly comparable to the DeepSeek moment, because the underlying logic is the same. At the end of the day, cutting-edge models make people can't help but ask: why invest in them? What makes them so important? Compared to building something truly usable, are they even a moat or competitive advantage? Is there a cheaper, better path? This reinforces a view I've held for a while: a large number of future agentized tasks won't necessarily require cutting-edge models. At least I don't think so. Now that there are powerful, far cheaper, open-source models, companies will very likely choose them.
Putting Sino-US tech relations aside for now, I don't want to call it a tech war. But over the past few months, the discussions I've participated in have all shifted to model interoperability: which model should be used for which task? Now people are realizing even more: since I have so many other options, why do I need Fable? Why do I need GPT-5.6? Models will increasingly become commodities, and the competition will lie in how you apply them, how you design workflows, and who controls the data. I don't want to overstate this as a total upheaval, but the narrative that "cutting-edge models are an enormous moat" is definitely becoming less tenable.
But the problem is, Kimi isn't drastically cheaper than other models, nor has it truly beaten the latest cutting-edge models on benchmarks. For example, it hasn't really outperformed Fable on most benchmarks. It beat OpenAI's 5.5 and 5.6, and Opus 4.8, but not the top-tier Fable model. So it's not distinctly better. It's also not drastically cheaper than Grok 4.5 or Meta's Muse 1.1 — the latter two are cheaper, and on some benchmarks can also compete with the Opus 4.8 and GPT-5.5 series models.
To me, that last tiny gap between DeepSWE and cutting-edge benchmarks is far less important than whether it roughly reaches the 4.8 or 5.5 level. Reaching that level qualifies you to compete. So from that perspective, it's very important that it can beat Fable on some benchmarks.
Host: Speaking of Grok 5.5 — I guess they're still sprinting on that version.
Ranjan Roy: They (xAI) definitely are. Grok 4.5 is very strong on coding benchmarks, reportedly more token-efficient, and most importantly, cheap. Over the past eight days, we've seen Meta and Grok enter the fray. They didn't release models better than Opus 4.8, GPT-5.5, or GPT-5.6 — they released models that are almost as good, but cost only 25% or 50% of the price.
Host: So I think the most compelling part of this release is that it's not absurdly cheap, but definitely cheaper: $3 per million input tokens, $15 per million output tokens. It's 40% cheaper than GPT-5.6, and 70% cheaper than Fable. Again, Meta has already made a strong entry, Grok too. More crucially, this is a Chinese company saying: "We're not seizing the market by undercutting prices drastically. We compete directly on quality, with a slightly lower price, and users have more control over the model."
So I think this is different. The core of DeepSeek was: can we build something that's not as good, but roughly sufficient, and far cheaper? Now, for the first time, we're seeing something that's truly competitive in actual quality. Is it better than Fable 5 or GPT-5.6, or worse? I think time will tell. But this isn't something absurdly cheap, or a throwaway bargain you'd buy on Temu. This is a genuinely high-quality, reasonably priced product. It makes people realize that cutting-edge capabilities may no longer be a moat.
As for whether this will completely rewrite the investment cycle — for example, if Anthropic stops investing once it realizes cutting-edge status no longer forms a moat — we can talk about that another time. But this is important, because from now on, people will increasingly ask: which model is the most cost-effective and efficient for a specific task? I'm already hearing that, and it will only become more common. Six months ago, people would say: why wouldn't you use a cutting-edge model? It's obviously the best.
Ranjan Roy: There are several important points here.
In the past, you might have thought: "I want to use the most powerful model."
So you'd give engineers Anthropic or OpenAI's cutting-edge models to build products for the company. Moonshot's model will have open weights. As long as your team is capable enough, you can download the weights and build applications on it yourself. Of course, the model is so large that self-hosting it requires massive infrastructure. Only institutions like governments or J.P. Morgan can afford that.
Apple has already done something similar with Gemini, right? The question is, now that Apple has made these advances, will it continue working with Google, OpenAI, or Anthropic to leverage a strong model like it did with Apple Intelligence, or will it switch to open-source models? On the other hand, models will be downloaded and deployed on various cloud platforms, and commoditization starts right here. You just listed Moonshot's quoted price — can cloud providers run it more efficiently and sell it to users at an even lower price?
That's going to be interesting. OpenAI and Anthropic's API business has always been: build a product far superior to any other on the market, and earn extremely high profits via premium services. Now two forces are crashing in: Meta and Grok are offering similarly performing but far cheaper models; Chinese open-source models are also delivering performance on par with their best models at competitive prices. People in the API business naturally have to ask: how much profit is left in this business? Will all models become commodities?
We also have to acknowledge that Anthropic has fully pivoted to the enterprise market, and is doing well. OpenAI doesn't have the same advantage in the enterprise market yet, but it's clearly investing heavily and taking plenty of actions.
Profit Shifting Layer by Layer: When the Model Layer Becomes the Terminator of Monopsony
Ranjan Roy: I think the Kimi K3 release hits OpenAI's positioning and pivot the hardest. Suddenly, the question of "why exactly is OpenAI important" is far more acute than it was a week ago. I want to read an analysis from Gavin Baker, investor and Managing Partner at Atreides Management. He's been on podcasts a lot lately, but this analysis is exceptionally good, clearly spelling out who gets hurt and who benefits.
He says: "Kimi K3 could be a major inflection point for AI. It's likely negative for Anthropic and OpenAI, but overall positive for almost every other company in the world. If only two or three dominant cutting-edge labs control 90% of inference margins, that's a net negative for every other layer except those two or three labs. Those two or three labs will of course benefit, but they will become monopsony buyers of power, data centers, semiconductors, and hyperscale cloud providers, and obviously will vertically integrate into all those layers over time, while completely swallowing the application software layer. Anything that reduces margins in the model layer and increases competition there benefits every other layer of AI: power, semiconductors, hyperscale cloud providers, new cloud providers, and of course software."
I think that's a very accurate summary.
Host: "Monopsony" is one of my favorite terms. I remember learning it in undergraduate economics: it refers to a dominant buyer, rather than a dominant seller. I'm dying to see Anthropic's S-1 filing. I'm sure they can always tell a coherent story. But again, is that story built on 90% inference margins? Train, pour money in, stay ahead in models, then reap huge profits. Kimi K3 is directly challenging that logic. I think this opens up more opportunities across the entire AI story, and allows more players to enter the field.
The sentiment on Twitter shifted so quickly. Six to eight months ago, everyone was saying Claude and Anthropic were unstoppable, the greatest thing in the world; now everyone is obsessed with which model is best for which task. Kimi K3 has stirred all these factors together. Competition is a good thing, it's exciting.
Ranjan Roy: Baker also says that the computing power required to run open-source models is exactly the same as that needed to run closed cutting-edge models of similar size and architecture. Per token, Kimi K3 is roughly comparable in scale to GPT-5.6 Terra. If anything, that suggests its computational efficiency might be lower. But if the model layer earns a little less, every segment of the infrastructure layer earns more; and it's huge good news for software companies. This scenario could be brought about by open-source cutting-edge models like K3, or by vertically integrated cutting-edge model companies like Meta, SpaceX, or Google. That's what we're talking about. Google is entering the fray too.
Baker believes that either outcome will compress margins in the model layer. Vertically integrated model companies don't particularly care which layer the profits end up in. That's why Google, by keeping pace with OpenAI and Anthropic on model capabilities back then, made things so uncomfortable for the latter two. That's also why Grok 4.5 and Muse 1.1 are just as important as Kimi K3. And all these events converged in this week.
Host: So what do you think OpenAI and Anthropic should do? OpenAI can still build speakers, hardware, cloud services, or launch other new business lines. But what about Anthropic? It's completely reliant on this for survival, and inference margins are its lifeblood. Their pressure is clearly much higher now. If you were one of them, what would you do?
Ranjan Roy: Let me read a bit more of Baker's analysis. He also addressed this question, then I'll share my own thoughts. He says: "Kimi K3 is only potentially