HomeArticle

The non-speaking AI has a valuation of 1.4 billion.

融资中国2026-09-28 09:20
The viral JEV

In September, the AI industry has never lacked buzz, with new models from major tech giants taking the spotlight one after another. Following past conventions, the center of public discussion should have revolved around these new releases. Yet over the past week, what developers have been flooding feeds with on X, GitHub and various tech groups is a model that refuses to even speak in complete sentences.

It is named Jev, launched only on September 15. Just a few days after going online, its official API was overwhelmed by swarms of developers and became inaccessible. By September 21, the company simply removed the waitlist to open access to all users, granting each registered user a $5 credit, equivalent to roughly 120 million tokens. It is rare this year that a new model "unleashes traffic" in such a manner.

The company behind Jev is TypeSafe AI, headquartered in San Francisco. The firm stayed under the radar for two years, and emerged with two things upon its debut: one is the model, the other is a $40 million seed round led by DCVC, equivalent to around 280 million RMB. Forbes cited a person familiar with the matter saying that the valuation of this round is $200 million, or roughly 1.4 billion RMB.

What makes people more curious is its founder Diogo Almeida. He spent four years at OpenAI, participating in the R&D of RLHF, InstructGPT, ChatGPT and GPT-4. RLHF refers to Reinforcement Learning from Human Feedback, a technology that plays an indispensable role in enabling ChatGPT to understand instructions and respond appropriately. In other words, he is one of the people who once taught large language models to "speak human language". After leaving OpenAI, the product he built turned out to be the exact opposite.

A "Mute" Model That Has Gone Viral Among Developers

Most people who have used large models for classification have encountered such troubles. You only want it to judge whether an email is urgent and reply with "yes" or "no", but it insists on spitting out a long analysis first, which is slow and consumes a lot of tokens. Jev's solution is simple and straightforward: it simply prevents the model from generating natural language. It takes in messy status information, and outputs decisions of pre-defined types set by developers, with each answer attached to a calibrated probability. TypeSafe named it the System One decision model, a reference to the "fast thinking" concept in the book *Thinking, Fast and Slow*. It is designed to handle tasks that humans can make a call on at a glance.

Almeida explained this design in a podcast. Jev only has three types of outputs: Choice picks one from given options, Score is used for sorting or comparing with thresholds, and Noul answers yes-or-no questions but returns a probability, leaving it to the program to decide whether to enter a certain branch. To put it plainly, it is built to be read by code.

Speed and low cost are the direct reasons for its skyrocketing popularity. The official stated end-to-end latency ranges from 70 to 500 milliseconds, which is 40 to 200 times faster than cutting-edge large models. It charges $0.042 per million input tokens, and outputs are completely free.

TypeSafe has also released more striking data. In its own workflow tests, the speed of some tasks can be increased by up to 193.6 times, and the cost can be reduced by up to 444.6 times at the slowest. However, these data are obtained from the company's own internal tests and have not been verified by independent laboratories.

The company also notes that actual performance varies depending on tasks and comparison methods. Third-party data is more conservative: after independent testing, it is found that its classification consistency is comparable to DeepSeek V4.1 Flash, with a median latency of 0.32 seconds. Even if we discount this performance, the price is still attractive enough, and some articles even compare its pricing to the "Mixue Ice City" in the AI industry.

What really pushed its popularity to new heights are all kinds of creative use cases from developers. The official set the first example, letting Jev play the original *Doom* game. It makes about ten decisions every second to control the character to move, dodge enemies and shoot, and running continuously for one hour costs roughly $7.

The Browser Use team integrated it into browser automation. The whole process of opening a webpage and pulling up flights from Zurich to London only takes 7.1 seconds, costing $0.0039. Real-time communication company Ably made a more intuitive comparison: within a few minutes, Jev made 47 decisions, while Gemini, Claude and GPT only made one or two. Chinese developers also did not lag behind: when using ordinary large models to play card games, obvious lag would occur, but after switching to Jev, a single decision only takes 0.7 seconds, and before the card-playing animation finishes, the next round has already been dispatched. All these scenarios require fast and intensive decision-making, but each individual decision is not difficult.

These playful use cases are fun, but the moves made in the business world are even more illustrative. Vercel integrated its AI gateway with Jev immediately after its launch. Developers use it as a referee for large models, completing the audit of whether a large model's behavior is compliant and fact-based within tens of milliseconds, and outputting a factual consistency score. When the cost of computation becomes so low that it can be ignored, things that were impossible to compute before suddenly become feasible.

A Key Contributor to ChatGPT Threw a Cold Water on the Industry

Almeida's resume is highly reputable in the industry. In the list of GPT-4 contributors, OpenAI listed him under the section of "Foundational RLHF and InstructGPT Work". He spent four years at OpenAI polishing ChatGPT's responses, and left in 2024 to start training a new foundational model. Among the other two co-founders, COO Sasha Sheng was previously a research engineer at Meta and FAIR, and Erik Gafni serves as CTO. The three of them worked in obscurity for two years, and almost no one outside the circle had heard of the company.

Normally, a person with such a background starting a business would most likely build a more conversational model. But Almeida did the opposite: he criticized exactly the methodology that he personally helped build.

His logic is not complicated. The RLHF approach collects human preferences and optimizes based on them, which essentially puts humans directly into the training loop. Its goal is to satisfy humans, rather than let software operate autonomously. He put it more bluntly to Forbes: the industry has long been optimizing for humans, and models are already better than humans at the skill of pleasing people. He once shared a joke in a speech: someone sent ChatGPT a fart sound effect and asked it to honestly evaluate this "music" he created. ChatGPT replied that it created a very eerie atmosphere. This kind of smooth response is harmless when a human is supervising. But if you want software to run unattended in the background, a model that is clearly wrong but always tries to act like it is right will become a huge problem.

So he dares to say that the next era will not be the era of Claude Code. He said he also likes to use Claude Code, but it still belongs to the "assistance era", which is fundamentally based on RLHF. Coming from one of the creators of ChatGPT, this statement is somewhat ironic.

To a large extent, investors are betting on this judgment and the person who made it. The comment from American investor Trace Cohen is very representative: in his view, giving a $40 million seed round to a company that has not yet delivered products is betting on the team's background, not the product. Researchers who left OpenAI have always been the most sought-after entrepreneurs in Silicon Valley, and this kind of valuation logic is nothing new. Interestingly, TypeSafe itself was not very confident about its product before launch.

Almeida later recalled in the podcast that the feedback before launch was quite bad. Non-technical colleagues worried that what they were selling was "vitamins" rather than "painkillers". The company had almost no revenue, more than half of the testers could not understand the product, and the first reaction of those who did understand was to ask how to get purchase approval. He was also very scared, so he pushed hard to launch the product as soon as possible.

He probably did not expect the situation after launch. Applications for raising call limits flooded in from all sides. According to his statement in the podcast, the daily call volume has exceeded 1 trillion tokens, and continuous calls even late at night indicate that the program is running in the background, rather than being tried out by humans and then abandoned. This figure is currently only claimed by the founder, but the detail of late-night calls is indeed more illustrative of whether the product is actually being used than the number of registered users.

The name also has special meaning. Jev is taken from the economist William Stanley Jevons, the scholar who proposed the Jevons paradox. In the 19th century, steam engines became more and more coal-efficient, but Britain's coal consumption kept rising instead. The intention of this name is not hard to guess.

The "Great Unbundling of Intelligence": Whose Interests Are Being Disrupted?

Jaya Gupta, a partner at Foundation Capital, wrote a long article putting Jev's popularity into a broader context, which she called the "Great Unbundling of Intelligence". Her observation is that the miracle of large models in the past three years lies in "bundling": information extraction, search, sorting, judgment, tool selection, response generation and even advanced reasoning are all handed over to the same sufficiently capable model. Developers saved effort, but the hidden costs were overlooked. An agent will split one human goal into hundreds or thousands of machine decisions, and tiny differences in cost and latency at each step accumulate, ultimately determining the economic model of the entire product. She calls the waste the "decoding tax": the software only needs a single "approve" or "block" decision, but the large model has to first generate a piece of text, and the program then translates the text back into a decision.

This cost will eventually be borne by large model companies. Gupta mentioned in the article that according to reports, Anthropic's gross margin exceeds 80%, and a simple classification task, just because it is called via Claude, has to be charged at the price of a cutting-edge model.

Her deduction is that low-cost, predictable high-frequency operations will be gradually unbundled, leaving cutting-edge models with more ambiguous, longer-cycle, harder-to-verify problems. Each call may be more valuable, but the number of calls will decrease. The work at the application layer will also change accordingly: developers need to split tasks, and purchase the sufficient, cheapest intelligence for each step. In the past, a system might only sample 1% of behaviors for inspection. After judgment becomes cheap enough, it becomes possible to inspect all of them: customer service conversations, transaction records, and contract clauses can all be checked one by one. The scenarios Almeida is targeting also fall into this category, such as the "dark data" that large companies have stockpiled for years but never analyzed because calling large models was too expensive. He cited the example of insurance underwriting: after the model reads the documents, it directly outputs the probability that a certain property has experienced a fire.

The story sounds promising, but the problem of moat has also emerged almost simultaneously. Just two days after Jev's release, at least six replica projects popped up on GitHub. Some people used a small Qwen model to train on a laptop for less than two hours, and built a decision model with the same API.

Many of these replicas are designed to be low-cost alternatives. The fastest-growing project Laya got 18,000 Stars in 5 days. Kev built by Jared Palmer is based on Qwen3.5, fully compatible with TypeSafe's official SDK, so business code does not need to be modified, and users can switch to local deployment just by changing the request endpoint. Nimble from Bespoke Labs is based on Qwen3.5-9B, with a test accuracy of 90.1%, close to Jev's officially announced 93.2%. Amid the buzz, Alibaba's Tongyi Qwen unexpectedly became the base model for this wave of replicas. Replicas also have shortcomings: some evaluations show that for scenarios with arbitrary queries and dozens of long options, Jev still performs better for now.

Jev's own limitations are also obvious. It is a closed-source API, users cannot get the model weights, and data has to be transmitted to TypeSafe's servers, which is a barrier for many enterprise clients. The official claim that "it does not hallucinate" also needs to be viewed with caution: it only guarantees that the output will not go beyond the format defined by developers, not that the answer is necessarily correct.

Almeida does not shy away from these issues. He admits that calibration is not perfect, and the model can make many mistakes, but as long as the benefit is high enough and the threshold is set appropriately, it can still be put into use. A host found after trial that it performs very well on single-step judgment, but the performance gradually declines when there are too many steps. His response is that which tasks are suitable for System 1 ultimately depends on actual performance. External pressure is also mounting: major tech giants are continuously driving down inference costs and increasing speed. How long the advantages claimed by TypeSafe can be maintained remains to be tested by the market.

In the past few years, the industry has been competing for whose model is better at thinking. The buzz around Jev this week made the industry sit down seriously for the first time to calculate another account: how much does it actually cost to make one single thought?

This article is from WeChat Official Account "Rongzhong Finance" (ID: thecapital), written by Wang Tao, edited by Wu Ren, and published with authorization from 36Kr.