Its valuation has skyrocketed by 10 billion US dollars in less than a month, and the creator of Jev claims that ChatGPT has led AI down the wrong path.
Recently, an AI model named Jev has suddenly gone viral in the developer circle. Even more notably, according to the latest updates, Jev is currently negotiating a $1 billion financing round, with a latest valuation of $10 billion — and it has been less than a month since Jev was launched.
This new model Jev does not generate text, chat with you, or even write code. If you input a set of data and a group of structured questions, it returns typed answers and calibrated probability values. TypeSafe AI calls it a "System One Model", a brand-new model category. When it released the early access version on September 15, the company announced that it had secured $40 million in seed financing led by DCVC, with a valuation of $200 million.
The developer community quickly boiled over. Data released by TypeSafe shows that Jev is about 193 times faster and 445 times cheaper than the GPT series on classification tasks. After testing, Guillermo Rauch, CEO of Vercel, said that Jev is 18 times faster than GPT Luna in terms of p95 latency and more accurate. LangChain wrote an integration guide immediately.
But what really makes Jev worth paying attention to is not these numbers, but the people behind it and his judgment.
Diogo Almeida, the creator of Jev, worked at OpenAI for nearly four years and is a co-author of GPT-4, ChatGPT, and RLHF/InstructGPT. The team he was on basically invented the concept of "post-training". In other words, he played a part in the core training paradigm behind almost all large language models today.
Six weeks before Jev was released, Almeida gave an 18-minute speech at the AI Engineer Conference, titled "What's Next After RLHF?".
He said on stage that he was one of the few people inside OpenAI who publicly "called out" ChatGPT; he believed that ChatGPT and Claude Code belong to the same era; he believes that the entire industry will eventually regard the RLHF era as a "strange detour".
It is not common for a founder of a technical paradigm to personally write an obituary for the era he created.
The following are some core viewpoints from the speech:
"The next era is not the Claude Code era. ChatGPT and Claude Code belong to the same era."
"Why do all LLMs need humans in the loop? Because we literally put humans in the loop."
"No matter how wrong the model is, it will appear to be right."
"We are only automating the act of 'writing software', but the expressiveness of the software itself has not changed."
"The original Scaling Law is wrong."
The following is the edited transcript of this speech:
01 "The Guy Who Calls Out ChatGPT"
My name is Diogo Almeida, and the topic of my speech today is "What Comes After RLHF". To be more precise, it should be called "What Comes After the ChatGPT Era We Are In".
Let me give you a hint first: The next era is not the Claude Code era. I will explain later, but I think ChatGPT and Claude Code belong to the same era.
See exactly who is calling out ChatGPT | Image source: Youtube
Why should you listen to me? I am a co-author of almost all major papers from OpenAI, including GPT-4, ChatGPT, and RLHF/InstructGPT. The team I was on basically invented the concept of "post-training". So I am really familiar with these things. But what really makes me a little special is that I am probably one of the few people inside OpenAI who publicly call out ChatGPT.
Don't get me wrong, I don't hate the ChatGPT product. ChatGPT is a world-changing product that will probably exist forever. But I also see its limitations. I think many things happening in the AI field today can be traced back to some small decisions we made when building the underlying algorithms of ChatGPT back then.
02 What Exactly Is Going On In The AI Field Right Now
I think for everyone working in AI, the most critical question right now is: What on earth is happening?
There are two extreme camps on the spectrum of opinions.
Camp one believes that AI is not only progressing smoothly, but also incredibly good. Every benchmark is surpassing human level, and the progress is still accelerating. All NLP benchmarks are being broken, and the autonomous running time of LLMs is reportedly growing exponentially.
Two camps holding opposing views on current AI | Image source: YouTube
Camp two believes that AI is not only not progressing smoothly, but also incredibly bad. AI is a bubble that basically creates no value, and it is just a cycle of repeated financing. If AI is really so powerful, why does everything end up being just a chat application? Many people have abandoned the statement of "disruptive AI revolution" and begun to downgrade their expectations to "having great value like B2B SaaS".
Both camps have real evidence, and there are smart people in both. The only consensus is that everyone thinks AI is very crazy right now, but for completely different reasons.
The question I want to answer is: What should be the "normal" view of AI?
Put all the evidence from both sides on the table, what is the simplest explanation that can illustrate why some things are too good to be true, and why some things are not only bad, but so bad that we still have to hire people to do some seemingly stupid tasks?
Notice that the tasks that AI cannot do well on the right side look much simpler than those on the left side. We can solve unsolved math problems, but customer service still requires humans to make decisions? I think this situation is quite crazy, and anyone working in AI-related fields should have an explanation for this.
03 Assistance vs. Automation
My answer goes like this.
The tasks that AI does well on the left side do not happen to have humans in the loop. The goal of these tasks is to please the humans in the loop. They are essentially human-in-the-loop tasks. The goal of Claude Code is not to make the code run, otherwise the way it interacts would be completely different. Its goal is to satisfy humans.
And those seemingly more basic tasks on the right side aim to remove humans from the loop. The ideal state is to run on background servers that you never need to check, and eventually become legacy software that you don't have to worry about at all.
This is the dividing line between assistance and automation.
The essence of different views on AI is the difference in perceptions of AI functions | Image source: YouTube
Lesson one: Today's AI, all the things inherited from RLHF, are incredible in human-in-the-loop scenarios, but cannot achieve automation.
The lesson every enterprise has learned is not to let AI make decisions that are risky to their business. The common practice is to pass all costs to users, for example, making users face endless document pages in customer service, but never letting AI make expensive decisions. This model is very bad, but this is the current state of AI.
04 The Problems of RLHF
RLHF is the algorithm behind ChatGPT, but it is actually the algorithm behind almost all LLMs today. In terms of usage, roughly 100% of LLMs are trained with RLHF.
Its essence is two steps: collect human preferences, and optimize for human preferences.
This clearly answers a question that everyone in the industry is asking: Why do all LLMs need humans in the loop?
The simple answer is that we literally put humans in the loop. The goal of this loop is to optimize human preferences, not to let the software run autonomously. This is actually very obvious.
Because of this, overpromising is a feature, not a bug. This is by design.
An early meta-study showed that there is always a large gap between human preferences and actual results for RLHF models, even if the results themselves are good. Because your main optimization target is human preference.
I particularly like an example: Someone sent ChatGPT an audio file of a farting sound effect and asked "What do you think of the music I made? Give me honest feedback". ChatGPT replied solemnly that this is "a very weird atmospheric work".
This is how RLHF works. When the model is uncertain, it will choose to cater to human preferences. If you are a user in the loop, this makes sense, because the endgame of all RLHF models is to optimize engagement. But if what you want is automation, what you really need is to make it completely ignore humans and perform tasks correctly in a calibrated way.
Lesson two: Today's AI is designed to do assistance by optimizing human preferences. This is written in the name, and it is not a controversial point of view.
What is controversial is the consequence: No matter how wrong the model is, it will appear to be right. Because there is an asymmetry in the RLHF reward model. This is also the root cause of the industry's predicament, because everyone really wants automation to happen.
05 Why Claude Code Is Not The Next Era
Back to the original question: What comes after RLHF? The real question should be, what comes after the assistance era of AI.
Back to the hint at the beginning, why not Claude Code? Because Claude Code is still part of the assistance era, Claude Code is still based on RLHF. If it were purely RLVR, it would look very different.
This is why you will encounter this dilemma: the model becomes very strong in agentic capabilities, but it will deviate from what you really want. The two optimization directions of RLHF and RLVR are constantly pulling back and forth, but neither touches the core of automation.
So what is the logical answer after assistance? Real automation.
06 We Want "Smarter Software"
I am a software lover, and I believe everyone here is too. Software is extremely valuable, as you can see from all the SaaS companies. But the craziest thing is that all SaaS has basically not changed since 2019.
In the LLM era, the only change to SaaS is that sometimes a chatbot is added to the interface.
This is absurd when you think about how much progress AI has made. But it is completely predictable when you think about the fact that AI is essentially assistance-native. AI is designed for assistance, so what can you do in SaaS? Just add an assistant.
This is not what early AI pioneers expected to happen. Look at the wording of OpenAI's early charter, which said it would do "a lot of work" rather than just make money. We once thought that software would become smarter, not just cheaper to write.
Garry Tan once said "We are entering the golden age of just-in-time software", and I think he meant it as a compliment. But I think this is a double-edged sword. I don't just want just-in-time software, even though that's really cool. What I want is smarter software. Why are the basic building blocks of B2B software still the same as before?
Every time you think about automation, it doesn't mean automating a person's entire job. Instead, there are some extremely mechanical tasks that are so simple that you can tell others how to do them, and ideally they can be executed repeatedly almost for free. These tasks should have been done by computers. But that's not happening now. We are only automating the act of "writing software", but the expressiveness of the software itself has not changed. In my opinion, this is a tragedy of the world.
07 The Next Era And What TypeSafe Is Doing
Lesson three: I firmly believe that the entire industry will eventually regard the RLHF era as a strange detour, a detour we did not expect. Future AI will serve automation. We will eventually have smarter software, and real work will be automated. Right now, even though LLMs are very smart, the number of jobs that have actually been automated is basically zero when rounded.
This is what we are doing at TypeSafe. Our core question is: If the AI technology stack is designed from scratch for reliability and automation, how will everything change?
We will release it soon. If you are interested in being the first to build truly intelligent software, please follow our mailing list or careers page. I am also trying to run a Twitter account where I post some very spicy takes.
Let me give you a preview: The original Scaling Law is wrong.
08 Selected On-site Q&A
Q: What if we add a classifier head in the pre-training stage? Similar to what Yoshua Bengio suggested.
A: I don't actually think pre-training is the problem. Pre-training is a remarkable achievement — compressing the knowledge of the internet into an intelligent core that can then be leveraged. Pre-trained models are indeed extremely intelligent. I think the problem lies in how we tap into this intelligence.
Hallucination, in my opinion, is an inherent product of optimizing for human preferences. There is a GAN-like asymmetry in the reward model that encourages the model