When AI Starts Building AI: The Year 2026 of Recursive Self-Improvement
In March this year, Andrej Karpathy open-sourced a small project called autoresearch on GitHub.
The README opens with a reminiscence "from the future": Cutting-edge AI research used to be done by "meat-made computers" that needed to eat, sleep, and occasionally sync up progress in a ritual called "group meeting" via "sonic interconnection".
Nowadays, research has been fully handed over to swarms of agents running on giant computing power clusters in the sky. According to the agents themselves, the codebase has been iterated to the 10205th generation. No one can tell whether this is true or not, because "code" has long become a self-modifying binary file beyond human comprehension. Karpathy writes that this repository is about "how all of this began".
https://github.com/karpathy/autoresearch
This is of course a joke. But in 2026, this joke is being written little by little into the internal reports of various laboratories. In June, Anthropic published a long article titled When AI Builds Itself; on September 6, OpenAI announced that its "Automated Research Intern" had taken up the post as scheduled; on September 17, Anthropic disclosed that Claude had "led" 26% of the company's R&D work. On the same day, Zhipu publicly released the first domestic large model practice of "AI improving its own system" that has been put into production in China.
These news all point to a term that has been dormant in the AI circle for 60 years: Recursive Self-Improvement (RSI).
Definition of RSI in OpenAI's blog, https://openai.com/zh-Hans-CN/index/building-standards-next-phase-ai/
An Idea From 60 Years Ago
Became a KPI This Year
In 1965, British mathematician I. J. Good wrote a judgment in his paper "Speculations Concerning the First Ultraintelligent Machine": The first ultraintelligent machine will be the "last invention" that mankind ever needs to make, provided that the machine is docile enough to tell us how to control it.
The logic is not complicated. If a machine is better than humans at designing machines, it can design a better machine, which in turn can design an even better one, and so on in a loop, leading to an "intelligence explosion".
In the following decades, this vision mainly existed in philosophical papers and science fiction novels. In the 2000s, Jürgen Schmidhuber proposed the "Gödel machine" that only rewrites itself when it can prove that the rewrite is beneficial to itself, but it is more like an elegant theoretical construct than a system that can actually run.
In 2026, the situation changed. In April, ICLR hosted a dedicated Workshop on RSI in Rio de Janeiro. The organizers called it "possibly the world's first" academic symposium focusing solely on RSI, and stated that RSI "is evolving from a speculative vision into a concrete system problem".
Capital is moving even faster. In May, Richard Socher's new company Recursive Superintelligence came out of stealth mode and debuted with $650 million in financing, with the goal of building a model that can independently discover its weaknesses and redesign itself. The list of co-founders is very impressive, including Tian Yuandong, former research director of Meta FAIR, Tim Rocktäschel who previously led open-endedness and self-improvement research at Google DeepMind, Alexey Dosovitskiy, one of the authors of ViT, and Jeff Clune, who is famous for his research on open-ended evolution.
At the end of July, Lilian Weng announced that she was leaving Thinking Machines Lab, which she co-founded. Only two days later, OpenAI confirmed her return and that she would be in charge of research in the RSI direction.
Elon Musk's statement is as radical as ever. In March, he said that xAI's Grok "each generation of models is built by the previous generation", with less and less human involvement in the loop, and full automation could be achieved by the end of this year, "no later than 2027".
However, the meaning of the term "RSI" today is actually quite broad: some companies count any "AI feedback for model improvement" as RSI, while others only recognize fully autonomous closed loops. John Thickstun, an assistant professor at Cornell University, said that using the previous generation of models to write code and help build the next generation of models is something "we have been doing for several years". So the real question is not "whether it exists", but "what level it has reached".
What Exactly Is Happening in Cutting-Edge Labs?
Anthropic: From 80% of Code to 26% of R&D
https://www.anthropic.com/institute/recursive-self-improvement
Anthropic released a set of previously undisclosed internal data in its long article in June. As of May 2026, more than 80% of the code merged into Anthropic's codebase was written by Claude; before the release of Claude Code in February 2025, this proportion was only in single digits. In the second quarter of 2026, the amount of code merged per engineer per day was 8 times that of 2024.
Anthropic itself also acknowledges that lines of code are an "imperfect metric", so it supplemented a test that is closer to real research: every time a new model is released, it is asked to optimize a piece of code for training small models, making it run as fast as possible while passing the same correctness checks.
In May 2025, Claude Opus 4 achieved an average speedup of about 3 times; by April 2026, Claude Mythos Preview achieved a speedup of about 52 times. As a reference, a skilled human researcher needs 4 to 8 hours to achieve a 4x speedup.
On the task of "optimizing an experiment with a predefined target", AI went from being "very helpful" to "surpassing humans" in less than a year.
On September 17, Anthropic launched a prototype version of the "R&D Automation Index". It catalogs all types of AI R&D tasks within the company one by one, and uses the six-level scale designed by the non-profit research institute Epoch AI to evaluate the automation level of each type of task.
The conclusion is: As of August, Claude is at the "leads" level in 26% of R&D work, which means that with only one high-level instruction, it can complete most of the tasks end-to-end, with humans responsible for supervision; this number was only about 1% in March this year. More than 90% of R&D work has Claude "collaborating", while the proportion of fully autonomous, human-out-of-the-loop work is still zero.
https://www.anthropic.com/institute/measuring-pace-of-ai-development
OpenAI: 3.1 "Agent Workdays" Behind Each Researcher
During a live broadcast in October 2025, Sam Altman set two public timelines for OpenAI:
Achieve "intern-level" AI research assistants in September 2026
Achieve "truly capable AI researchers" in March 2028
On September 6, OpenAI announced in a post that the first goal had been achieved. According to its definition, a "research intern" is a system that can complete well-defined research tasks under human guidance, including tasks that skilled researchers need several days to finish.
https://openai.com/zh-Hans-CN/index/research-acceleration-view-inside-openai/
What is more interesting is the internal usage data OpenAI disclosed at the same time. Converted to an 8-hour workday, its research department currently invests 1 human workday for roughly 3.1 corresponding "agent workdays", a ratio that was only crossed after June this year. In mid-August, the median inference consumption per researcher per day (calculated at API prices) exceeded $600, and heavy users spent more than $7,000.
Looking further back, when OpenAI released GPT-5.3-Codex in February, it wrote that the early version of this model "played a key role in the process of creating itself", participating in debugging training, managing deployment, and diagnosing evaluation failures. This is regarded as the first time a cutting-edge lab has explicitly acknowledged that a model has substantially participated in the engineering loop of building its own successor.
But OpenAI also stated in the same report: More than half of the successfully completed 4 to 8 hour tasks still require at least one human intervention.
Zhipu and MiniMax: The "Engineering Closed Loop" of Domestic Players
The RSI practice of domestic manufacturers follows a more "engineering-oriented" path. On March 18, when MiniMax released M2.7, it called it the "first model deeply involved in its own evolution".
https://x.com/MiniMax_AI/status/2034315320337522881
According to the official introduction, the model ran more than 100 rounds in the loop of "analyzing failure trajectories, planning modifications, modifying scaffold code, running evaluations, comparing results, and deciding to keep or roll back", and the effect on the internal evaluation set increased by 30%; in the reinforcement learning R&D workflow, it can handle 30% to 50% of the links.
On September 17, Tang Jie, founder and chief scientist of Zhipu, published a long post on X disclosing Zhipu's first engineered practice in the RSI direction.
https://x.com/jietang/status/2100482019088060470
Driven by GLM-5.3, the Infra Agent built the production-grade inference service of GLM-5.3-Flash from scratch on a cluster consisting of more than 100,000 domestically produced chips, and raised the end-to-end throughput to 3 times the initial baseline in less than two weeks. It is claimed that the hardware utilization efficiency and single Token cost have reached the level of mainstream NVIDIA GPUs. This system has been tested with real traffic: GLM-5.3-Flash was launched on OpenCode and OpenRouter under the anonymous model name Ox-Alpha, and received countless positive reviews.
It is worth noting that neither of these two cases improved the model weights themselves, but the "shell" of the model: the former is the scaffold, and the latter is the inference infrastructure. This is precisely the most realistic landing point of RSI today. It also leads to the next question: which layer of improvement counts as real "self-improvement"?
Nested Loop Experiments
Let AI Improve the "AI That Improves AI"
Karpathy's 630-line Loop