He personally created o1 and o3, but suddenly announced: Mankind can retire forever.
Hey, have you heard about this?
The good days for human researchers are only two years at most in total.
Two years from now, when humans conduct AI research, it will be just like how humans play chess today.
You can still play it for sure, but no one will care at all how well you perform.
This is the latest bold claim put forward by Jerry Tworek, the former lead of OpenAI's reasoning model team.
He joined OpenAI in 2019 and stayed there for seven years. Back when reinforcement learning could not be scaled up no matter how hard people tried, he stuck to the task and pushed it forward relentlessly. The two landmark models o1 and o3 were both developed under his personal leadership.
"We want to build AGI." That was the entire roadmap Ilya presented at the all-hands meeting when he first joined OpenAI in 2019. Seven years later, he left to start his own business
What's more striking is that he is not the only one saying this. The most popular internal meme among AI researchers right now goes like this —
We only have a few more days of work left to do, so hurry up and get it done while we still can, then we can all retire and rest together.
Everyone is passing this around as black humor. But everyone who tells the joke knows perfectly well in their heart.
Most likely, this joke will become reality.
And this is just the appetizer. Throughout the interview, he kept throwing out such blunt remarks one after another:
- Agents do have creativity, but it is an extremely low-quality, massively stacked kind of creativity. They generate all kinds of ideas, which are generally very poor
- There are probably only 30 to 50 people in the world who have truly mastered the full process from training to deployment of a cutting-edge model, and everyone else is assisting them
- The assumption that Transformer is the optimal solution is almost impossible to hold
- During the seven years I spent at OpenAI, there were only three to four serious attempts to replace the underlying architecture
- The premise that the world has to rely on us to work to keep running is not valid in the first place
The following is the organized version, Enjoy.
Half of the work is already done by non-human workers
The two-year timeline has nothing to do with benchmark scores, nor with computing power.
He is counting how many parts of the research work are still being done by humans.
And this work has long been split into two parts.
One part is coming up with ideas, figuring out the right direction to move forward. The other part is executing the work, turning ideas into runnable code and then retrieving the data.
And the execution part has been largely taken over by Agents.
Tworek founded a company called Core Automation this April, whose slogan is to build the world's most automated AI lab.
At their company, the full cycle of one experiment has been directly compressed from one month to one day. The efficiency has been increased by 30 times!
However, for the idea generation part, Agents cannot take over for now. The reason is —
They are absolutely species with high creativity.
But they are squandering creativity in an extremely low-quality, massively stacked way. Their ideas are indeed varied, but they are also generally very poor.
There are only 30 to 50 people in the whole world
Sounds like the idea generation part is still stable for now.
But there are not many people left who can actually come up with valid ideas.
During the interview, the host revealed that an internal OpenAI researcher told him privately —
There are probably only 30 to 50 people in the world who can truly understand the full end-to-end process of training and deploying a cutting-edge model.
Everyone else is assisting these dozens of top talents, including the vast majority of full-time employees in that top-tier company.
Tworek did not refute this point.
He also thinks that any top team operates in this way. A small number of people set the course, followed by a whole roaring execution machine.
To see how sought-after these dozens of people are, just look at his own recruitment list.
Rohan Anil, co-founder of Core Automation, comes from Anthropic, and previously worked at Google DeepMind.
Anmol Gulati, who worked on Gemini at DeepMind, was also recruited. Even Julia Villagra, the former head of human resources at OpenAI, joined the company.
All labs are poaching talents from the same pool. And this pool only has dozens of top talents in total.
So the two-year timeline has nothing to do with all of humanity.
It refers to how long these dozens of people can keep their current roles.
The architecture has only been seriously modified three or four times in seven years
Since there are only dozens of top talents left, what is the last wall standing in front of AI?
Tworek's answer is only one word: Transformer.
In his opinion, this architecture is almost impossible to be the optimal solution.
First of all, the model cannot continue learning after deployment. No matter how much you talk to it, it will not get stronger, and it will remain the same in the next conversation. The context window also cannot support continuous learning. He said that after using Codex for more than 20 minutes, he has to compress the context once.
Secondly, if you try to make up for this shortcoming through continuous fine-tuning, you will find that it is not only extremely inefficient, but also leads to catastrophic forgetting: the model will forget old knowledge after learning new ones.
So most of the work the entire industry has done around Transformer in recent years is essentially making it cheaper, and very few people are actually making it more capable.
What Tworek really wants is not actually the architecture itself.
What he wants is a model that can continue learning at test time, a model that can keep growing from user interactions and user data. Replacing the architecture is just a means to achieve that goal.
No one in the industry is unaware of these principles. But after so many years, why is Transformer still firmly in place?
The reason behind it is so simple that it seems ridiculous — hardly anyone has really tried to replace it.
During his seven years at OpenAI, there were only three to four serious attempts to replace the underlying architecture.
The process goes like this.
Researchers first write a small-scale verification experiment, which takes at least three months to run. Only when the results look promising will they dare to scale up the experiment.
And so-called scaling up means you have to convince about ten people who hold the power of life and death over the project, and then these ten people will invest three to six months into your bottomless pit of a project.
In the end, your project will either be crushed by the huge accumulated momentum of Transformer, or partially absorbed and turned into a tiny component of Transformer.
Seven years, three or four attempts. This is the total output of the world's top AI lab in terms of architectural innovation.
"Jare, use these computing resources"
He himself once ran into exactly this dead end at OpenAI. And what broke the deadlock was just one sentence.
At that time, those dying experiments showed a little "sign of life". It was not good enough, but there was finally a glimmer of hope.
At this critical moment, Jakub Pachocki, who later became the chief scientist of OpenAI, found him.
"Jare, all these GPUs are for you, see if you can make the results you are working on bigger and more powerful."
"Now you have all these GPUs in your hands." When he repeated this sentence, o1 didn't even have a name yet
This sentence pulled him out of a paradox —
You must first get results before you are eligible to apply for computing resources.
But you clearly need to have that computing resource first to get the results.
Tworek said that most cutting-edge research directions are trapped in this paradox. The fierce hidden battles in labs for competing for computing resources are essentially all looking for a way out of this paradox.
And the exit is sometimes just a little bit of confidence from the leadership.
As long as someone in a leadership position says, I hope you can get a decent computing resource quota to do this research to the fullest. That's enough.
Once the gate is opened, the momentum cannot be stopped later.
Reinforcement learning crossed several orders of magnitude starting from this node.
Then o1 was successfully developed, and the reasoning model path that countless people had once sentenced to death was brought back to life by him.
100,000 dollars is equivalent to one top expert
This time, he doesn't need to wait for anyone's approval anymore.
The most expensive part of the entire chain is translating abstract ideas into code that can actually run. And this is exactly the task that today's Agents do best.
For example, Core Automation used this set of methods to work on GPU kernels.
Specifically, they gave a QR decomposition kernel to a programming Agent to run for four weeks, spending about 100,000 dollars in calling costs, and finally increased the speed of this kernel to 60 times its original level.
This task belongs to low-level performance engineering, which focuses on how to make a section of matrix operation run faster on the graphics card. Normally, it requires very few top experts to manually adjust line by line.
And there are very few such experts in the whole world.
Now 100,000 dollars can get the same output as one such expert. The Agent does not need to be convinced, nor does it take three to six months.
The barrier that had been stuck for seven years was broken through in this way.
What else can we do in the future?
Even the most difficult architectural problem has been loosened, and the steering wheel in the hands of those dozens of people will not be held for much longer.
At the end of the interview, the host asked him this question clearly: when that day really comes, what else can humans do?
In response, Tworek described two scenarios.
The first one is ancient Greece.
People meet in the square, talk about philosophy slowly for a whole day, then go to exercise, eat olives and drink wine.
He laughed as soon as he finished describing this scene, admitting that it was probably just a projection of his own wish.
"We meet in the square and then chat." His original words describing the daily life of humans in the post-AGI era, he laughed right after saying that
The second scene is high school, or university.
He thinks people should retain the pursuit of "excellence" itself. Keep learning new things greedily, and push both the body and mind to the limit.
This is a bit like professional sports. There is no economic reason to force you to sweat and bleed, but the pride in human nature makes you want to reach for the thing called "greatness".
We have to find all kinds of ways to do this.
Because in the next world, there will no longer be such things that "if you don't do it, the world will collapse". The world will keep running on the infrastructure we have already built.