HomeArticle

When AI-written papers outperform over 99% of human-submitted manuscripts, Chen Yongchao raised four unanswered questions.

星连资本2026-09-04 10:34
Chen Yongchao shared the self-evolving large model, and ApexResearch can independently produce top conference papers.

On August 8, 2026, Chen Yongchao, Assistant Professor of the School of Artificial Intelligence, Tsinghua University, and Founder of HyperEvolution AI, shared his insights on self-evolving large models at the Link-X Demo Day 2026 event with the theme "From Static Replication to Autonomous Discovery: Self-Evolving Large Models Define the Next Generation of Super Intelligence".

What is Super Intelligence?

On August 8, Chen Yongchao, Founder of HyperEvolution AI, opened his speech with this question.

Interestingly, just a few days ago, Google's legendary leader Jeff Dean also announced his departure, co-founding Discovery Loop with Sanjay Ghemawat, Oriol Vinyals and Quoc Le to enter the RSI (Recursive Self-Improvement) track.

"This is almost exactly what we are working on," Chen Yongchao said. This is not a case of keeping up with others, but a signal that this field is moving from the academic fringe to industry consensus.

01

The True Intelligence:

It Is Not About What You Have Learned, But Whether You Can Have a "Eureka Moment"

François Chollet, former Google researcher and founder of Keras, mentioned in an article published in 2019 that there are flaws in the current methods we use to test AI intelligence. "What we are doing now is essentially looking at its capabilities on various benchmarks and what it has learned. But higher-order intelligence should refer to its ability to learn unknown things, that is, the ability to innovate," Chen Yongchao noted. This is consistent with the theme "Eureka Moment" of the first session of the conference. What we care about is not how much knowledge an AI has memorized, but whether it has the ability to generate unexpected breakthrough inspirations.

What current large models are doing is compressing and remembering the existing knowledge of the world. What HyperEvolution AI aims to do is to enable them to create knowledge that the world has never known before.

This track did not emerge today.

In 1965, British cryptographer and mathematician I.J. Good published the paper "Speculations Concerning the First Ultraintelligent Machine", in which he first proposed the concept of "Intelligence Explosion". He predicted that when artificial intelligence surpasses human intelligence and can improve itself autonomously, it will trigger an unstoppable exponential leap, and "the ultraintelligent machine will be the last invention that man need to make".

But this concept was impossible to realize at all before the emergence of large language models.

In 2023, the industry began to develop preliminary self-evolving systems, allowing large models to optimize prompts and agents on their own. Chen Yongchao was also working on related projects at that time. But he soon found the upper limit: after optimization reaches a certain level, the performance will no longer increase no matter how much adjustment is made.

In 2024, Japan's Sakana AI developed the "AI Scientist" that enables AI to write papers on its own, verifying the feasibility of this direction — but the output can only reach the workshop level, with a gap to the standard of top conference papers.

The turning point came in February 2026. After the release of Claude Opus 4.6, Chen Yongchao tried it himself and found that it could indeed make research more efficient and automated. Based on this judgment, he developed Apex Research.

02

The Singularity Has Arrived:

34 Papers, 2 of Which Surpass 99% of Human Submissions

Apex Research is a system that can conduct research autonomously without human intervention.

In the past few months, Apex Research has completed the entire process from literature research, method design, code implementation, experimental analysis, chart drawing, LaTeX paper writing, to self-review and iterative optimization, and can now produce top conference papers at the level of NeurIPS and ACL.

In May this year, Apex Research autonomously produced 34 papers and submitted them for ACL review. The results came out in July:

- The average initial review score of 11 papers is ≥3

- The average score of 2 of these papers reaches around 3.7

What does a score of 3.7 mean? Chen Yongchao took himself as a reference: "When I was a PhD student submitting papers to ACL, only one of my submissions got an average score of around 3.7, and all the rest were lower than 3.7."

99% of human submissions get scores lower than 3.7. That means the research capability of AI has now reached the level of human PhD holders.

Apex Research is not only used in the field of AI research.

Recently, the HyperEvolution AI team cooperated with Professor Zhang Jingzhao from Tsinghua University to conduct pure mathematical theoretical research using this system. It turns out that AI has improved a mainstream consensus in this field over the past decade. After repeated verification, mathematicians confirmed that this proof is correct, and it has now been written into a paper that is about to be submitted.

03

Four Questions Without Standard Answers

At the end of the speech, Chen Yongchao raised four questions that he is currently thinking about.

How to raise the innovation capability of AI scientists from the PhD level to the level of Newton and Einstein?

Chen Yongchao shared a view he called a "radical opinion": All research is essentially permutation and combination. Those groundbreaking researches, such as Transformer, only explored a permutation and combination with extremely low probability density, which happens to be correct. Most researches keep lingering in areas with high probability density, so they are not attractive enough.

"To improve the innovation capability of AI, we need to increase its thinking diversity, so that its thinking can drift to those permutation and combinations with extremely low probability. How to make this happen and how to judge whether an idea is feasible — this is why we are developing self-evolving models," Chen Yongchao said.

When AI can produce scientific discoveries in batches, how should we evaluate them?

If AI scientists reach the level of Einstein, they will automatically generate massive research results. But the evaluation of scientific conclusions in human society relies on time — many important discoveries can show their influence only after decades. In the future, the evaluation of science may be more difficult than the generation of science.

Will the paper as a medium still exist?

Papers have existed for hundreds of years. In the past, due to the limited information processing capability of human beings, researchers often spent one or two years on experiments, and finally compressed the results into a few pages of paper for dissemination. There is a certain amount of information loss in this process: loss occurs when researchers write papers, and readers supplement content with their imagination when reading, resulting in information asymmetry.

"But this is not a problem for AI. The information processing capability of AI is much stronger than that of humans. We can load all research processes and experimental settings into the system, and let the AI on the other side read it, so that it can accurately obtain all information and reproduce the results." Chen Yongchao's judgment is that for the research community with AI scientists as the main body, papers will likely no longer be the main dissemination medium.

Who owns the research achievements?

If AI can produce high-quality research, should the achievements belong to the person who uses the AI, or to the AI itself? Who should take the corresponding responsibility? Chen Yongchao gave an example: a high school student using this system may produce a large number of high-quality papers. If he directly submits the papers claiming that he wrote them himself, it will be very difficult to identify the truth. The existing social evaluation system may be broken.

"Now research is dominated by people with high IQ. But in the future, research may be dominated by people who have more computing power, or more capital."

04

What HyperEvolution AI Wants to Do:

Unlock Undiscovered Discovery

There are no standard answers to these four questions.

Chen Yongchao said that a company developing self-evolving models should not wait until all problems are figured out before taking action. Instead, it should turn these problems into discussible propositions one by one in the process of practice.

"HyperEvolution" means super derivation. When AI can not only answer questions, but also raise questions, design experiments and verify hypotheses, the definition of research is being rewritten. What HyperEvolution AI wants to do is to make this change happen faster, safer and more reliable.

The following is the full text of the speech.

Hello everyone, I am Chen Yongchao, Founder of HyperEvolution AI. Thank you very much for the invitation from the organizer today. The theme of my sharing is: From Static Replication to Autonomous Discovery, Self-Evolving Large Models Define the Next Generation of Super Intelligence.

Recently, the self-evolution track is very popular. Two days ago, Jeff Dean from Google also publicly announced that he is working on this direction. I checked the direction he is working on, and it is almost exactly the same as what we are doing.

What is Super Intelligence?

 

First of all, I want to discuss a question with you: What is super intelligence? Or what is artificial superintelligence?

I want to quote an article written in 2019 by François Chollet, former Google researcher and founder of Keras, who I really like. The main point of this article is that the current methods we use to test the intelligence of AI systems are flawed.

To put it bluntly, what we are doing now is looking at its capabilities on various benchmarks and what it has learned. But he thinks this kind of intelligence is relatively elementary. Higher-order intelligence should refer to the ability to learn unknown things — that is, its learning capability. And this capability is essentially the innovation capability.

This is very similar to the theme of today's forum "Eureka Moment" — we need to see if an AI system has the ability to generate unexpected breakthrough inspirations.

Current large models are compressing or remembering the existing knowledge of the world. What we want to do is to enable them to create knowledge that the world has never known before. This is the AI Scientist, or the self-evolving large model, which is what our company (HyperEvolution AI) is currently working on.

This Track Did Not Emerge Today

 

In fact, this concept is not new. In 1965, British cryptographer I.J. Good proposed the concept of "Intelligence Explosion" — allowing machines to design better machines, iterate generation after generation, so that the intelligence level becomes higher and higher, and finally reach the effect of intelligence explosion. But this concept was impossible to realize at all in the past, until the emergence of large language models.

In 2023, people began to develop a very preliminary self-evolving system — allowing large models to design their own prompts, or design their own agent systems. At that time, I also had a related research work. But later we found that the upper limit of this method is relatively low. Because after optimization reaches a certain level, no matter how you optimize its prompts or agents, the performance will no longer improve.

In 2024, Sakana AI from Japan proposed the AI Scientist. They let AI conduct AI research autonomously and write papers. But later they also found that the papers written by their AI Scientist at that time could only reach the workshop level.

Last May, I interned at Google Research and DeepMind, and I was working on the AI Scientist track exactly. They wanted me to build a better system at that time. But I didn't do it immediately, the reason is very simple — I thought the large models had not reached that level of capability at that time, and the research results produced would be relatively trivial.

The turning point came in February this year. After Claude Opus 4.6 was released, someone told me that the research results produced by this model are already very impressive. I didn't believe it at first, but after I tried it myself, I found that it can indeed make research very efficient and automated. Based on this judgment, we built a system called Apex Research.

34 Papers, 2 of Which Surpass 99% of Human Submissions

 

This system can already conduct research autonomously without human intervention, and produce papers at top conference level — such as papers for NeurIPS or ACL.

In May this year, Apex Research autonomously produced 34 papers, which were submitted to the ACL review of this year.

The results came out in July:

Among the 34 papers, 11 papers got an average initial review score of ≥3

The average score of about 2 papers among them is around 3.7

What does a score of 3.7 mean? When I was a PhD student submitting papers to ACL, only one of my submissions got an average score of 3.7, and all the rest were lower than 3.7. 99% of human submissions get scores lower than 3.7.

That means the research capability of AI systems has now reached the level of human PhD holders in many cases. If your AI system can reach the level of a PhD from Tsinghua University, it will be very valuable.

It is not limited to the AI field. We cooperated with Professor Zhang Jingzhao from Tsinghua University to conduct pure mathematical theoretical research using this system. It turns out that Apex Research can already discover very powerful conclusions — it has produced a proof that improved a mainstream consensus in this field over the past decade. Our mathematicians have read it repeatedly for several times, and basically confirmed that it is correct. It has now been written into a paper that is about to be submitted. This tells us one thing: AI doing research is very promising. In the future, most researches may be completed with the help of AI.

Why Can't Current Large Models Achieve This?

 

But later we also found a problem: after optimization reaches a certain level, if you want to go further, you can't achieve the goal only by optimizing prompts and harness, you have to solve it through model training. This is what HyperEvolution AI is doing now — training the next generation of self-evolving models.

On the other hand, we are currently in the agent stage. But we will definitely transition from the agent stage to the next stage — we need to enable the model to explore unknown knowledge of the world. This is also why the self-evolution track suddenly became so popular.

But there is a key problem here: the existing first-generation models are likely not to reach the goal of self-evolution. The reason is very simple — the innovation capability and the stability-seeking capability are contradictory in many cases.

It is very difficult for an agent to have both extremely strong innovation capability and be extremely stable and error-free. You can think about it, human society is the same — a person with very strong innovation capability often has unconventional ideas, and is easy to put forward