HomeArticle

Why is Google betting big on RSI?

深流研究所2026-10-09 15:38
The much-debated RSI: what is the new controversy all about?

"Our model has officially surpassed me. I have nothing left to teach it."

On September 30, Zirui Wang, a researcher at Google DeepMind, wrote on social media.

On the same day, Google released Gemini 4 Argon, the first model in the Gemini 4 series. This is Google's most important flagship release of the year.

This half-joking exclamation soon led to a bolder speculation: Has Google achieved RSI?

RSI, short for Recursive Self-Improvement, refers to AI that participates in improving itself or developing the next generation of AI.

Today, many companies have included RSI in their research goals and investment narratives. Relevant papers, evaluations and products are also emerging intensively.

Google is one of the companies that has invested the most and taken the most aggressive actions.

In September alone, Google successively announced four related advances, covering model R&D, search strategies, agent workflows and more.

Behind this intensive layout is Google's sense of urgency in the large model competition.

In less than a year, Gemini has fallen from the top of the rankings to being teased by some users as "American Doubao".

The flagship model was delayed, core talents left, and key directions such as programming agents were first defined by competitors. The risks Google is facing have evolved from falling behind in a single generation of models to gradually losing control of the industry's development rhythm.

RSI seems to be a lifeline that Google must seize.

According to a report by Business Insider, Google co-founder Sergey Brin is not satisfied with the progress of Gemini, and has pushed the team to devote more energy to RSI.

Faced with the doubt that it is "falling behind in AI", why does Brin pin his hopes on RSI?

What is the ongoing debate over the much-discussed RSI really about?

RSI seemingly means "AI improves AI", but there is no unified definition in the industry so far.

The debate mainly revolves around three questions:

What exactly can AI improve? Can one improvement reduce the cost of the next progress? How to prove that the improvement is truly effective?

First of all, what AI improves is not necessarily only the model itself.

Agents can rewrite prompts, adjust tools, optimize training code, and also redesign experimental processes.

Anthropic regards code agents as the pre-stage of RSI. Claude first becomes a R&D assistant, and then gradually takes over program operation, agent invocation and long-cycle tasks.

In Anthropic's view, the loop is not truly closed until the model can more autonomously complete training, evaluation and subsequent system development.

What OpenAI focuses on is to what extent research work can be automated.

On September 6, OpenAI announced that it had reached the stage goal of "Automated Research Intern": the system can, under human guidance, complete tasks with clear boundaries that originally took skilled researchers several days to finish. Its next goal is to evolve into an "Automated AI Researcher" before March 2028.

Google's understanding is more inclined to systems engineering.

As long as AI can continuously improve training code, algorithms, toolchains, chip design and data center scheduling, and use these gains to shorten the development cycle of the next generation of models, recursive effects may appear in the entire R&D system.

Secondly, repeated iteration does not equal recursion.

AI can continuously write code and run tests. If the capability ceiling and working method of each round do not change, it is just repeating labor at a faster speed.

The real recursive effect requires that one success can improve the efficiency of subsequent search, experiment or R&D.

For example, training kernel acceleration can shorten the training time of the next-generation model; data centers releasing more computing power can support more experiments; optimized agent workflows can allow the same model to undertake more complex tasks.

These improvements are scattered in software, hardware and organizational processes. As Yao Shunyu, a researcher at Google DeepMind, said: "RSI is a continuous spectrum, not a moment that arrives suddenly."

Last but not least, the most difficult question is how to verify the improvements made by AI.

There are relatively clear standards for whether the code can be compiled, how much the kernel speeds up, and how much memory is saved.

It is often more difficult to judge whether a model has acquired capabilities that can be transferred to new tasks, and whether a set of automatically generated training methods is truly better than the original solution.

Recently, RSI Bench launched by Scale AI, AI4AI-Bench focusing on algorithm design, and Φ-Bench covering GPU kernel, training, inference and service fields, are all testing the pre-required capabilities for RSI.

The improvement object, degree of autonomy and verification mechanism constitute the core debate about RSI today.

The answers from different companies also correspond to different technical routes.

From models to data centers, how does Google lay out RSI?

Google's investment in RSI was originally scattered in multiple links such as models, agents, search strategies, chips and data centers.

In the past month, these relatively independent technical routes have begun to converge.

Gemini 3.8 Flash, released on September 2, sent a direct signal.

Google stated that long-running agent loops participated in the development of this model. In complex tasks, agents repeatedly plan, execute, check and repair, assisting the team in completing model evaluation and improvement.

Yao Shunyu commented that this is "a small step" for model capabilities, but a "giant leap" for RSI.

The next step is to reduce ineffective trial and error in R&D.

On September 14, researchers from Google and two American universities released Dream-RSI.

In order to find faster GPU programs, code agents need to explore multiple routes, most of which will fail. If computing power is allocated evenly according to a fixed strategy, the system will easily consume resources repeatedly in hopeless directions.

Dream-RSI is equivalent to equipping agents with a continuously updated "exploration map".

It records the code, error reports, running time and final score of each attempt, and then judges from historical experience: which direction is worth continuing, and which direction should be abandoned in time.

When AI faces similar tasks, it can avoid detours and allocate more computing power to more promising directions.

A week later on September 21, the Google Cloud AI Research team, together with three American universities, released RRSI to further transform the working mode of agents.

Whether an agent can complete complex tasks depends to a certain extent on the external Harness.

How to write prompts, which tools to call, how to split tasks, and how multiple agents cooperate will all affect the final result.

In the past, this "console" was mainly manually adjusted by engineers based on failure cases.

RRSI allows AI to participate in modifying its own console: modify prompts according to task feedback, add or remove tool calls, re-split tasks, or adjust the division of labor among multiple agents.

Dream-RSI solves the problem of "where to go", and RRSI solves the problem of "how to get the work done well".

They do not modify the weights of the base model, but enable the same model to use experience, tools and computing resources more efficiently.

This is exactly Google's current idea of promoting RSI:

Start from the periphery of the model, accumulate improvements in links that are easy to measure and verify, and then bring the gains back to model R&D.

AlphaEvolve, which was released earlier, has already brought this idea into the production system.

Its operating logic is similar to an automated engineering team. Gemini puts forward candidate solutions in batches, the evaluator tests them one by one, the system retains the versions with better performance, and then continues to explore around these versions.

At present, AlphaEvolve has been deployed in Google's data centers, chips and AI training processes, used to improve algorithms, optimize training, and find more efficient scheduling rules for computing clusters.

On July 10 this year, Google officially launched AlphaEvolve for commercial use on Google Cloud. Automatic optimization has expanded from an internal R&D tool to an external platform capability.

So far, Google's originally independent technical modules have been reorganized.

Models, agents and infrastructure are beginning to connect with each other, and the technologies and resources scattered within the company are incorporated into the same iterative loop.

A R&D production line for RSI is taking shape.

Why does Google have to bet on RSI?

For Google, RSI is not only a technical route, a logic of capital return, but also an opportunity to regain control of the industry's development rhythm.

In early August, Jasjeet Sekhon, Chief Strategy Officer of Google DeepMind, said at the Berkeley Agentic AI Summit that RSI is one of the core investment logics behind the current huge AI investment.

In 2026, Alphabet raised its capital expenditure guidance to 195 to 205 billion US dollars, and most of the funds will be used for servers, data centers and network equipment.

Google must prove that these resources can not only serve core businesses such as search and advertising, but also continuously improve its own R&D efficiency.

RSI just provides a kind of compound interest:

Computing power expands R&D capabilities, stronger models complete more experiments, experiments improve software and infrastructure, and the upgraded infrastructure then supports the next generation of models.

Google has a complete technology stack spanning software, hardware and application scenarios, and also has a large number of real business scenarios that can be automatically verified.

In the past, these business systems were Google's scale advantages. In the agent era, they may also become test sites for model training and verifying RSI capabilities.

In addition, RSI also provides Google with an opportunity to "overtake on a curve".

The AI competition in the past few years has proved that technological accumulation will not automatically translate into product leadership.

Transformer was born in Google, but ChatGPT first defined the mass entry point for generative AI.

Google has long studied reinforcement learning and reasoning, but OpenAI earlier shaped reasoning models into an independent product category.

Google has a huge code base and developer ecosystem, while Anthropic took the lead in establishing product awareness of programming agents with Claude Code.

Google has always been at the forefront of technology, but rarely holds the industry rhythm for a long time.

RSI may not allow a certain generation of Gemini to suddenly establish an overwhelming advantage. But it can turn one-time technological progress into accumulated R&D compound interest.

Nowadays, almost all leading AI companies are looking for their own compound interest engines, just taking different paths.

Anthropic's advantage is that its software engineering chain is shorter.

Claude Code is not only a product, but also continuously generates real task trajectories. Users ask Claude to read code libraries, modify files, run tests, and handle error reports. These processes constantly expose the capability boundaries of the model and feed back to training and agent design.

OpenAI's strength is that it has established a shorter product decision-making chain around cutting-edge models.

Research progress can be quickly integrated into ChatGPT, Codex and APIs. It does not have a complete chip and data center closed loop like Google, but it can more centrally allocate resources around the "Automated AI Researcher".

Google's weakness is its large organization and complex chains. But RSI may also turn this weakness into an advantage.

Once AI can take over more experiments, engineering and coordination work, Google's full-stack system, which used to seem cumbersome, may become a longer, more complete, and more difficult-to-replicate feedback loop.

According to a Reuters report in August, Brin is pushing more resources to RSI and urging core AI employees to fully commit to Gemini.

This means that RSI is no longer just a research topic within Google. It is substantially affecting resource allocation, product priorities and R&D rhythm.

What Google wants to catch up with is not only the most powerful model at present, but also the ability to continuously build the most powerful models.

Of course, today's RSI is still far from the "intelligence explosion" envisioned by I. J. Good.

At this stage, AI can only find better solutions within the scope delineated by humans. There are no mature answers to how to evaluate open research, how to prevent the system from exploiting rule loopholes, and whether local optimization can truly improve model capabilities.

But AI has already entered the process of developing AI, which will gradually change the competitive rhythm of the entire industry.

Beyond models, computing power and data, a set of AI-driven R&D systems that can continuously accelerate is becoming the core asset of leading companies.

Model performance determines the ranking for a while, and iteration speed determines how long the leading position can last.

This article is from the WeChat official account "Deep Flow Institute", author: Wu Jiangfeng, published with authorization from 36Kr.