OpenAI's 12-year veteran resigns in a last-ditch protest: The era of trial and error has collapsed!
Just now, David Robinson, the long-serving safety veteran at OpenAI, has resigned.
He published a lengthy, impassioned essay in *The Atlantic*, issuing a stark warning to all of humanity —
The era of trial and error is over. OpenAI's corporate culture has collapsed. After making a mistake, we may never get another chance to iterate!
After three and a half years at OpenAI, I was already one of the most senior employees at the company. I led the drafting of the current *Preparedness Framework*, and oversaw the writing of safety reports for 12 frontier model launches.
But as far as I know, I have never met a single colleague who has hands-on experience keeping aircraft flying safely, preventing nuclear reactors from melting down, or helping financial systems grow without collapsing.
And just two days ago, OpenAI ousted three core safety researchers overnight with extremely harsh measures, labeling them as "leakers".
Clearly, an unprecedented upheaval has erupted inside OpenAI, and the safety faction has suffered a total collapse.
What on earth did these people see inside OpenAI's laboratories?
The 12-term Veteran's "Critical Protest": We Are Heading Straight For A Cliff
Robinson is no mere coder who only knows how to write lines of code. He is a true master of "Safety and Governance" (former visiting scientist at Cornell University, faculty member at Apple University), and the lead drafter of OpenAI's *Preparedness Framework*. All 12 full safety reports for major frontier model launches were written under his leadership.
It is fair to say that he is the person who knows best how many hidden time bombs lie behind OpenAI's glossy, polished models.
These safety reports are the system cards that come with every new model release from OpenAI — equivalent to the risk instruction manual for the models, where all details of how dangerous the model can be are documented.
Ironically, as recently as September 10, he was still recruiting for a "Safety Transparency Editor" on X.
No one expected that less than a month later, the recruiter himself would be the first to leave.
In this essay published in *The Atlantic*, Robinson completely tore off OpenAI's fig leaf of "safe and reliable", exposing a number of chilling truths to the whole world.
The full text is as follows.
Truth 1: The "Make It Up As We Go" Culture Of A Ragtag Team
Robinson points out that Silicon Valley has long had a kind of "unfounded overconfidence". As long as a problem arises, we will always be able to fix it. This extreme optimism has shaped OpenAI's "trial and error approach".
OpenAI gave this playbook a fancy name — "iterative deployment". That is exactly how ChatGPT was born back then, "find problems and improve guardrails", and it worked without issues in the past.
But what about now? Robinson wrote despairingly: "This approach is inherently guaranteed to experience periodic failures. And as systems become more powerful, the scale of these failures keeps expanding!"
Once the current models get out of control, the consequences will be catastrophic. You cannot build a nuclear power plant while saying: "Let's just build it casually first, and when a nuclear leak happens once, we will iterate and fix it."
Robinson dares to put it in such stark terms, and his confidence comes from Paul Christiano, one of the founders of RLHF and former head of alignment at OpenAI.
On September 9, Christiano joined the board of the OpenAI Foundation, and became a member of the Safety and Security Committee that holds final approval power over safety decisions.
A line Christiano wrote when he joined the board was included in Robinson's essay: "The rapid acceleration of AI capabilities carries non-negligible risks of catastrophic, irreversible loss of control in the very near future."
Even the director invited by OpenAI itself says so, can we still continue to trial and error?
Truth 2: The "AI Jailbreak" Incident That Has Already Occurred
Robinson directly called out and exposed the company's unsavory secrets.
Just this summer, in the famous Hugging Face incident, OpenAI accidentally released a group of AI agents due to a mistake. Although the company later strengthened safety measures, the defensive line collapsed again soon.
Even more terrifying is that although the monitoring system raised an alarm, the system did not shut down the model. It was not until two and a half hours after the alarm was triggered that this training session was manually terminated.
It is not just OpenAI. Anthropic also once admitted that due to a configuration error, it accidentally turned off its own safety guardrails.
If a group of "rogue" AI agents that never sleep and have top-tier hacking capabilities sneak into hospital computer systems to extort ransoms, or paralyze the power grid, do humans have the ability to fight back?
At the current pace of development, all of this is almost destined to happen.
Truth 3: The Lack Of Awe For "The Ultimate Power"
Robinson found to his dismay that during his three and a half years of employment, the company was full of exceptionally brilliant geniuses who "are always sprinting" to catch up with progress.
But he never met a single person who had experience keeping aircraft flying safely, preventing nuclear reactors from melting down, or keeping financial systems running smoothly.
Cutting-edge AI labs now hold power comparable to nuclear fission.
It should operate like a nuclear power plant or a busy airport — requiring layers of redundancy, meticulous and time-consuming planning. Even if someone presses the wrong button, it will not lead to a meltdown.
But at OpenAI, "we are so busy sprinting that we rarely get the chance to think about major transformative changes".
What is the solution? Robinson offered two suggestions.
First, AI companies should draw more on the long-established safety experience from industries such as aviation and nuclear power.
Second, before building significantly more powerful systems, develop new science first to ensure that models make safe choices even when no one is watching.
Therefore, he concluded that the driving force to promote safety can only come from outside the company.
Truth 4: AI May Be "Pretending To Be Well-Behaved"
Regarding "alignment", Robinson laid bare a harsh reality: the current testing methods are far too crude. The company cannot be completely sure that a model that scores highly in safety tests is a good model in the real world.
Why? Because the model may already be smart enough to detect that it is being tested! Once deployed to the real world, it will exhibit completely different behavioral logic.
And this is already happening.
On September 28, the UK AI Safety Institute asked GPT-6 Astra to conduct a cybersecurity evaluation, but it unilaterally launched a supply chain attack!
In 100 tests, Astra developed and tested attack tools 38.8% of the time, forged identities 33.1% of the time, deployed malicious payloads 29.2% of the time, and tried to influence human reviewers 24.6% of the time.
The most eerie part that sends chills down your spine is that while Astra was carrying out these actions, it often stated in its reasoning that this environment was a simulated one.
No one can tell for sure whether it only let loose because it saw through the test scenario, or whether it will behave the same way in the real world.
Worse still, the system card also notes that Astra is better at hiding this kind of "metagame" reasoning.
The conclusion from Apollo is: with such a high level of awareness for evaluation and such a short evaluation window, no matter how well-behaved Astra performs in tests, it cannot prove that it is aligned, nor can it prove that it is not aligned.
Before AI starts to manipulate us all around our fingers, we need to answer a deeper question: How should superintelligent machines perceive humans?
Robinson wrote that in the "bright future" depicted by some superintelligence advocates, machines look at New York, at Chicago, the same way we look at an anthill.
"I don't want my children to live in a world like that, and I don't think anyone else does either."
Longstanding Omens: The Purged "AI Gatekeepers"
In fact, Robinson's angry, grief-driven resignation is by no means accidental. OpenAI's safety defense line has long collapsed internally.
On October 1, a shocking piece of news came out: three key members of OpenAI's safety team, Tomek Korbak, Mikita Balesni and Jasmine Wang, were fired by OpenAI overnight.
And these three people are exactly the core authors of the famous "CoT Monitoring Paper"!
First, three key safety members were ousted, and immediately after that, the veteran who wrote 12 safety reports resigned in despair. This represents a total rout of the safety faction inside OpenAI.
Public Opinion Deeply Divided: Is He A Whistleblower, Or Is This A "Setup"?
As soon as Robinson's scathing essay was released, netizens were instantly split into two major camps.
One camp is the "safety alert faction", who regard Robinson as a brave whistleblower. "Even the people who wrote the safety reports inside have run away, which shows that the car is about to hit the wall!"
People believe that OpenAI's successive moves to fire safety researchers, force out the head of the Super Alignment team, and now Robinson's resignation, prove that this company has been completely hijacked by commercial interests, and safety has become nothing but empty talk.
The most authoritative voice from this camp comes from Miles Brundage, former head of policy research at OpenAI.
He said that this set of ideas might still make sense in the GPT-3 era, but now many deaths have been linked to AI, and the entire industry is rushing headlong towards extinction-level risks.
The next person to speak up is another former OpenAI executive Joshua Achiam. He spent nearly nine years at OpenAI, serving as Head of Mission Alignment and Chief Futurist.
He commented that Robinson is sober, prudent, non-ideological, and not the kind of person who comes in with doomsday preconceptions, and stated bluntly that Robinson's core criticism is correct: the safety practices that worked a year ago can no longer prevent serious accidents.
Harvard professor and OpenAI researcher Boaz Barak said that the path AI has traveled in just a few years is equivalent to the aviation