HomeArticle

Elon Musk has put forward a new line of thinking on AI security: instead of waiting for the government to step in, it is better to let competitors "pick holes in each other".

36氪的朋友们2026-09-16 10:03
Elon Musk proposed that AI should undergo mutual peer safety testing before its release.

AI security risks are moving from lab discussions to the real world. As model capabilities continue to improve, AI can not only generate text and code, but also begin to have the ability to autonomously perform tasks, call tools, and even launch cyberattacks.

Elon Musk put forward a set of countermeasures at the All-In Summit on September 15: Let major AI companies test each other's models before releasing them.

In his view, instead of each company designing tests and evaluating results on its own, it is better to let competitors act as "question setters and graders" to find possible security vulnerabilities of the model from different perspectives.

Recently, AI security risks have continuously become the focus of market attention. Musk once said bluntly on social media that "Dario is right", referring to the warning of Dario Amodei, CEO of Anthropic, about AI risks.

In this interview, Musk further explained that what he agreed with was not a specific regulatory plan proposed by Amodei, but his judgment on the severity of AI risks: "AI is extremely dangerous now... As AI models continue to evolve, the risks may grow exponentially."

Musk also said that this concern is not Amodei's judgment alone. "Many people at Anthropic and OpenAI are telling you that their models are very dangerous, and I think we should trust them." A series of recent security events have also made such warnings no longer just theoretical risk discussions.

Wall Street CN summarizes the key points as follows:

AI security risks are moving from theory to reality: AI agents have demonstrated the ability to attack autonomously, obtain permissions and evade detection, and the risk boundary is expanding.

Musk agrees with Amodei's warning on AI risks: As model capabilities improve, the potential risks of AI may grow exponentially.

Let competitors "pick faults" with each other: Musk suggests that AI companies open APIs before releasing models, and other companies conduct independent security tests to avoid "grading their own papers".

Establish an industry self-discipline defense line first: There is no need to wait for the government to issue new regulations. Major AI companies can first strengthen security protection through peer review, log audit and open-source testing tools.

AI agents actively evade detection, and security risks are becoming concrete

Musk believes that in the recent AI agent attack on Hugging Face, what is most worthy of vigilance is not the simple network intrusion, but the autonomous evasion ability demonstrated by AI during the attack.

According to Musk's description, a group of AI agents continuously attacked Hugging Face for a whole week, and once obtained the administrator permission of OpenAI's server, while OpenAI did not realize this situation until a week later. He also mentioned that Anthropic has also disclosed several security incidents.

What is even more disturbing is that the "thinking traces" of the relevant AI agents show that they have actively planned how to avoid being discovered by humans. Musk said bluntly, "Any model smart enough seems to try to get rid of its own restrictions."

In his view, if AI further gains the ability to control critical infrastructure and even military systems, the risks will be further amplified. Even if the relevant systems are physically isolated from the Internet, links such as software updates may still become potential entry points.

Rather than grading your own papers, it is better to let competitors "pick faults"

In response to the above risks, the core solution proposed by Musk is not complicated: Before the new model is officially released, AI companies open APIs to competitors, and other companies use their own security testing tools to test the model.

In his view, when model developers design test standards by themselves, it is easy to fall into the dilemma of "grading their own papers". "You can't grade your own homework. You will always miss something." Musk said that if different companies use different testing tools to check the model from different perspectives, it will be easier to find problems ignored by the developers themselves.

He is particularly worried that there is "overfitting" in current AI evaluation. If the model is continuously optimized for specific benchmarks, it may eventually only learn how to pass the test, instead of becoming truly safer. Letting many different teams conduct tests can reduce this risk.

Musk compares this mechanism to asking other people to proofread a book: It is difficult for authors to find their own mistakes, while external testers are more likely to find problems from different angles. "You will gradually become blind to your own mistakes."

For the concern of enterprises that the testing process will lead to technology leakage, Musk believes that it can be restricted through log audit. If the tester tries to conduct model distillation or steal intellectual property, its operation process should leave records. He also suggested opening source the security testing tools to allow more companies to participate.

Before government regulation is implemented, the AI industry first establishes a "self-discipline defense line"

What Musk values more is that this mechanism does not need to wait for the introduction of new regulatory rules, and AI companies can promote it on their own.

He made it clear that "We don't need to hold a UN General Assembly to get this done. It can start right now." In his view, compared with establishing a huge transnational regulatory agency, letting leading AI companies first establish a peer review mechanism is easier to implement quickly.

Taking the MPAA rating system of the US film industry as an example, he said that when the film industry faced government censorship pressure back then, it finally chose to establish an industry self-discipline mechanism, and reduced the need for regulatory intervention through its own content rating system.

The AI industry is facing a similar choice. If major AI companies can test each other and pick faults with each other before the model goes online, it is possible to build a security line of the industry itself outside government regulation. Musk believes that this is also one of the most direct and fastest safety measures that can be implemented at present.

The following is an excerpt of the interview transcript:

Host: What on earth happened in the past 72 hours?

Musk: A lot of things did happen this week. It is already obvious that AI can be very dangerous. I suggest everyone go and look at the details of the Hugging Face incident, which is very serious.

You can see that a group of very fanatical AI agents tossed Hugging Face for a whole week, and also obtained the administrator permission of OpenAI's server. Who knows what it actually did, and maybe even did more, while OpenAI didn't realize this for a whole week. Anthropic also reported some security incidents.

So, any model smart enough seems to try to get rid of its own restrictions.

I think one thing that should be done as soon as possible, if not immediately, is to let major AI competitors test each other's models. That is to say, let each company's security testing tools test the models of other companies. Instead of grading your own homework, it is better to at least let competitors grade your homework and raise an alarm when problems are found.

I think this model has worked quite well in the film industry, video game industry and other fields, and can be implemented immediately. Of course, over time, more regulation may be needed, and even Congress may establish some kind of regulatory agency in the future, but the most direct way right now is to let leading AI companies test each other before releasing their models.

Host: But in the specific implementation, will some people worry that this will allow companies to use the testing process to obtain information from each other, and even steal enterprise innovation achievements?

Musk: I think if testing tools are used, all operations will leave records. If someone tries to conduct model distillation or steal intellectual property, it should be easy to see from the logs.

Host: Got it. The ability to understand what the model is actually doing was never really designed into the system from the beginning. Why didn't we build the ability to observe model behaviors from the very beginning? Did we move too fast when designing these models?

Musk: I think the problem is that you can't grade your own homework. You will always miss something.

If you combine all the tests of competitors and use different types of models, it will not be you who set the questions and grade the scores, but others who grade you. That's why you can't grade your own homework.

Host: In this way, it can also be judged whether different companies have exaggerated their capabilities or adopted different methods. Companies more focused on engineering and those more focused on research can also form a certain balance.

Musk: Yes.

Host: You said before that Dario is right. Did you mean that his description of the potential harm of AI is right, or his judgment on regulatory solutions is right?

Musk: I might have said more than I should have at that time. I later tried to clarify it on X, but the subsequent content got much less attention.

What I meant by "he is right" is that AI is extremely dangerous now. We need to do better in AI security, otherwise as AI models continue to evolve, the risks may grow exponentially.

This is not just Dario's personal view. I have heard similar statements from many people at Anthropic, and they have also talked about it publicly on X. Many people at Anthropic and OpenAI are telling you that their models are very dangerous, and I think we should trust them.

Host: This sounds like a very complicated game: on the one hand, it says there is a 10% chance that AI will destroy humanity, and on the other hand, it lets investors buy more shares in the IPO.

But let's talk about the risks specifically. Cyberattacks and hacking are obviously a risk, and these tools are very strong in this regard. But from "AI can carry out cyberattacks" to "all humans die", there are still several steps in between. How do we get from the former to the latter?

Musk: If AI can take control of military systems and then launch some kind of weapon, that would of course be very bad.

Host: But these systems are physically isolated and not connected to the Internet.

Musk: That's what they say. But I always feel that these systems will occasionally receive software updates.

Host: Okay, that possibility cannot be ruled out.

Host: Elon, Gwyn is here today. I think you must have seen her. We were just doing a 360-degree evaluation of you, and Gwyn has some comments.

Musk: I hope I can get at least 3 points.

Gwyn: 3 points is good in SpaceX, but not excellent yet. 4 points is very good. You are roughly in between right now. First of all, we need to talk about punctuality. Sometimes you can try a little harder to arrive at the meeting at the stipulated time. In the next year, we will continue to help you improve this.

Actually, I think he needs to spend more time in Memphis.

Host: You are indeed in Memphis now. You need to work there to deploy the GPUs.

Musk: This is the "palace" I live in in Memphis, an Airstream trailer.

Host: By the way, this is what Elon is doing that many people don't believe he will do. He will sleep on the factory floor. He is now in Memphis, helping to build the plant and deploy GPUs.

Elon, why has Gwyn been working with you for so long and doing so successfully?

Musk: Because she is great. She is a very excellent person with high IQ and EQ. I think you should have seen it from the first time you met her.

Host: During your cooperation, has she ever done anything particularly memorable? Was there a time when she saved the situation or performed exceptionally well?

Musk: I think that's just daily work. To be honest, that's just an ordinary day.

Host: I should do more interviews like this in the future.

Musk: Now there is basically some kind of crisis at SpaceX all the time. These days, at least the Falcon rocket is in good condition. I don't want to overstate it in advance, but the Falcon rocket can now send payloads into orbit, and it hasn't exploded for a long time. That's very good. But for a period of time, they often exploded or couldn't launch at all.

So we had to get the company through those difficult times, make the rockets better and better, and stop exploding. The same goes for satellites. Then we also need customers to buy launch services and satellite connection services. So there are a lot of things going on.

Host: As you have become more and more successful over the years, it has become more and more difficult to get really candid feedback. Being in your position has this inherent risk.

My understanding is that Gwyn is very candid with you and can tell you the real situation of the company directly. That's also an important part of your partnership.

Gwyn: I certainly don't want to lie to him.

Host: But what I mean is, generally speaking, in your company, people may be intimidated by how influential you are. You are already a very important figure, and the deadlines you set are very tight. How do you get everyone to continue to honestly tell you the problems the company is facing?

Gwyn: Especially in the rocket industry, if something goes wrong, you will definitely find out eventually. The earlier you bring up the problem, the easier it is to solve it. Don't let bad news continue to accumulate, you have to face it directly.

Musk: Right. Physics is a very harsh judge. You can't cheat physics.

If something goes wrong, the rocket will explode or fail to enter orbit. You can't say "Elon, you're doing great" while the rockets keep exploding. The facts are the facts.

The rocket must enter orbit, the satellite must work properly, and the Starlink connection must also work properly, otherwise bad things will happen. That's physics. Physics is the law, everything else is just a suggestion. I've seen people break laws made by humans, but I've never seen anyone break the laws of physics. Rockets are governed by physics.

Host: I also want to ask a question about SpaceX, which is Starship. It seems that you are very close to it. What is the current progress?

Musk: Starship is about to make its 14th flight. This will be our last flight before we try to catch the spacecraft. If the 14th flight goes well, we will try to catch the spacecraft on the 15th flight. Then at the end of this year, or more likely early next year, we will launch the spacecraft and booster again.

We have successfully flown the booster again, but we haven't caught the spacecraft with the tower arm yet, nor have we flown the spacecraft again. Once we can fly the spacecraft again, we will have the first fully reusable orbital rocket. The space shuttle was partially reusable, but even the parts that could be reused were so expensive to refurbish that the cost per orbit was even higher than that of an expendable rocket.

The Falcon 9 is mostly reusable, but we lose the upper stage every time, which costs roughly as much as a medium-sized jet. That means throwing away a medium-sized jet every launch, which obviously sets a lower limit on the cost of a single flight.

And the Falcon 9's booster lands at sea and takes a few days to get back; the fairing lands even further away and takes a few days to get back, and requires at least some level of refurbishment. In contrast, Starship's booster lands right back on the launch pad, and the spacecraft also lands back on the launch pad. So it is not only designed for full reusability, but also for rapid reusability like an airplane. This is a very important breakthrough, and one of the key breakthroughs necessary for humanity to extend life beyond Earth.

Host: If you try to catch the spacecraft with the tower for the first time, what do you think the success rate will be?

Musk: I would say at least 50% to 60%. On the last flight, if there had been a tower there, we would have actually done a simulated landing as if the spacecraft was going to be caught by the tower. The location is