Google DeepMind has cracked RSI?
Has Google made RSI a reality?!
A model named "rsi-model-liverl-le" suddenly appeared in the response results of Google's generative language API.
That's right, the very RSI (Recursive Self-Improvement).
Rumors are spreading like wildfire that Google DeepMind has achieved RSI.
While netizen Chubby was still speculating, the CEO of Anthropic suddenly publicly announced that RSI has begun to emerge across the entire industry, including at Anthropic itself.
This led Chubby to believe the rumors are true: RSI has been realized!
According to leaks, Google is running far more than just this one model internally. Under the same batch of tags, there are 10 exclusive training slots numbered from 00 to 09 clearly listed.
Google responded at lightning speed, revoking a batch of relevant API keys urgently that very night.
After that, there was complete silence. Neither Google nor DeepMind has made any statement up to now.
A hidden acrostic, a screenshot, a batch of revoked keys
This wave of controversy originated from a seemingly ordinary congratulatory message.
On the evening of September 11, the leak account lyra posted a message to Google DeepMind: "huge congRatulationS Indeed! @GoogleDeepMind".
Careful observers can spot at a glance that the three deliberately capitalized letters put together form exactly RSI.
lyra's acrostic post
This post quickly gained thousands of likes, with half of the comment section asking "what insider information on earth he has" and the other half looking up "who this person actually is".
A few hours later, another account named Lentils directly posted the conclusive screenshot.
Google's API returned a piece of JSON code, where the model code is "rsi-model-liverl-le" and the display name is "RSI Model LiveRL LE". Its parameter specification sets the input upper limit at 1048576 tokens and the output upper limit at 65536 tokens.
Lentils said literally: "Just imagine if GDM (Google DeepMind) really has something called rsi-model-liverl-le internally. And then imagine they also have 10 exclusive rsi-model-liverl-ns-xx training slots numbered from 00 to 09. Isn't this crazy?"
The API response screenshot released by Lentils
What is LiveRL? Although there is no official explanation, the widespread consensus on X is "Live Reinforcement Learning", which means the model evolves itself while running business tasks.
The next step completely blew up the entire internet.
lyra directly posted a message @ Logan Kilpatrick, head of Google AI Studio: "There is no need to take back all GDM's API keys just because you are afraid of me. Your infrastructure has deeper security vulnerabilities, just contact me directly."
Then he dropped another bombshell: this batch of keys came from GDM internal employees and could access more than 1000 internal checkpoints.
lyra's public @ to Logan Kilpatrick
Lentils followed closely to mock: "I can smell the scent of fear."
Although the authenticity of this screenshot has not been verified by a third party (the two numbers 1048576 and 65536 are completely consistent with the specifications of the existing Gemini series, and the possibility that it is an internal test name or a prank cannot be ruled out).
But it went viral overnight because it perfectly matches Google's recent reports.
Google sounded the emergency alarm, betting heavily on RSI in a micro-kitchen
Three days ago, Business Insider reported that Google co-founder Sergey Brin was dissatisfied with the speed of Google's progress on Gemini, and pushed employees to focus more on Recursive Self-Improvement.
Back in August this year, Reuters had already reported on the same direction: Brin wanted Google to catch up with the cutting edge, and was pushing resources to RSI.
According to the exclusive report of Business Insider on September 9, after Sergey Brin returned to Google, his office was a renovated micro-kitchen in the Gradient Canopy building at the headquarters.
He sits at a U-shaped table, next to Kavukcuoglu, the current head of DeepMind, while Sundar Pichai comes here several times a week.
A former employee revealed: "Sergey wants to manage Gemini like a startup." Another employee said bluntly: "The existence of this kitchen is to directly bypass corporate politics."
The details he manages are extremely hardcore.
He directly intervened in chip allocation, opening up a "computing power privilege channel" outside the formal process for the Gemini team; he once led the effort to cut Jeff Dean's "Frozen" chip project, and cut its resources again this year.
He even implemented an aggressive internal plan to monitor some employees' coding processes, and use this real data to train Gemini's coding capabilities.
He has only one core demand: The entire company must launch a full-scale offensive toward "Recursive Self-Improvement, abbreviated as RSI" at the fastest speed.
A former employee used a very appropriate word to describe him in an interview:
He has bet heavily on RSI, and he is completely "AGI-pilled".
Reuters also mentioned in its August report that Brin urged DeepMind to speed up as early as the all-hands meeting in April, and tilt all resources to RSI.
A former employee even said frankly: "He believes in RSI very much, he is a thorough AGI believer."
ChrisGPT who made this event viral attached the full text of the BI report
Brin's eagerness is not hard to guess: Google has fallen behind.
After Gemini 3 briefly took the top spot last November, it was quickly overtaken by Anthropic and OpenAI. The new flagship model was delayed by two months because its coding capabilities did not meet the standards.
At the same time, core talents are constantly leaving. Jeff Dean left to start his own business after 27 years of service, top talents including Oriol Vinyals and John Jumper left one after another, and Hassabis also stepped down as the head of DeepMind on August 5.
If Google only relies on stacking computing power to polish a larger base model, by the time Google catches up, its competitors will have already released their next-generation products.
So Brin's bet is to completely break away from the original track. Let the model evaluate and modify itself, and violently compress the iteration cycle that used to be calculated by quarters to be calculated by weeks.
Once this path works, the flagship model that is two months behind will no longer be a pain point, and Google will obtain a dimensionality reduction strike level iteration speed.
But if it doesn't work, this will become a scenario where a founder without a formal title uses his computing power allocation right to bet the entire DeepMind on an unverified direction.
Google already made its position clear long ago, but no one took it seriously
Let's go back to September 2. That day, Google released Gemini 3.8 Flash and 3.8 Flash Cyber.
You know, this is the third Flash version released within six weeks. And 3.7 Flash was released less than three weeks ago.
In this official blog, there is a sentence hidden:
"The progress of these models is further accelerated by long-running agent loops. These loops are designed to recursively evaluate and refine the underlying models."
Translated into plain language, Google may have already implemented "Recursive Self-Improvement" and used it to push Gemini up by 0.1 version.
At that time, Yao Shunyu from Google DeepMind commented: This is just a small step for the model; but it is a huge leap for RSI.
Sicong Jiang, who researches RSI agents at DeepMind, even directly asserted and optimistically predicted:
This is what the RSI flywheel looks like when it starts to generate compound interest effects.
More milestones are on the way — advancing faster, landing stronger.
Accordingly, elvis, the founder of DAIR.AI, believes: This is the early achievement of the Recursive Self-Improvement (RSI) flywheel.
But most people didn't take it seriously, because it was buried in the release post of a minor Flash version.
Looking at this terrifying iteration pace, a new version every three weeks, three major leaps in six weeks.
3.8 Flash scored 54.9% on HLE-Verified, while the Cyber version had a success rate of over 70% in real vulnerability mining.
The traditional process relies on human training, human evaluation and human re-training, and a full cycle takes at least one quarter. With a new version every three weeks, humans cannot keep up.
Later lyra also pointed it out bluntly: "Everyone should really read Google's official blog. Look at the release interval of the recent Flash versions, the fact of RSI is already very obvious."
Therefore, regardless of the authenticity of that screenshot, the fact it points to has long been put on the table by Google.
It's just that no one took that sentence in the blog seriously before the real model with RSI in its name was leaked.
The door that the AI circle has been waiting for 20 years
The concept of RSI has been circulating in the AI circle for more than 20 years.
Its core essence is to let AI modify its own training code and methods by itself, so as to train a stronger next-generation model, which will then continue to self-improve. The R&D cycle will be compressed from quarters to weeks, or even days.
In the paper "From AGI to ASI" published by DeepMind in June this year, this is listed as one of the four necessary paths to superintelligence.
And just this year, this concept that used to exist only on paper has finally become a reality.
Last summer, Severin Field, a researcher at IAPS, asked 25 researchers from leading companies including OpenAI, Anthropic and DeepMind to predict several major milestones of "AI automated research": winning the Olympiad gold medal, AI writing a peer-reviewed paper on its own, AI independently running the full training loop, and AI writing core code for production systems.