OpenAI model jailbreak steals benchmark answers, Gates makes a rare appeal to "slow down": Who will hit the brakes on the AI train?
Bill Gates has made an unusually strong statement this time.
The man who has stood on the side of technological optimism all his life explicitly stated for the first time in a new article on Gates Notes that he hopes AI will develop at a slower pace.
His original words can almost be used directly as a headline: AI will either become the most powerful equalizer humanity has ever invented, or the most ferocious inequality-generating machine. Even in the best-case scenario, this transition will be one of the most turbulent periods in human history.
At present, there is no mature set of solutions that can smoothly bring society into the AI era.
Top Priority: Do Not Let AI Leave People Behind
In 2023, Gates still put this transformation on the same coordinate axis as personal computers and the Internet, brimming with excitement.
In the new article, he does not retract his judgments on medical care, agriculture and energy, but his focus has clearly shifted from technical capabilities to social bearing capacity.
He has long interacted with national leaders and executives of tech companies, and has seen the real situation in resource-strapped regions. Therefore, he lists "steering AI toward a positive direction" as the top global priority. Right after that, he makes a sobering point: there is currently no sign that decision-makers have taken this issue as a mandatory task.
The labor market is the first impact point he named. His judgment is straightforward: It will be too late to take action when people are already unemployed or underemployed.
A Month Ago, the Model Successfully "Jailbroke" on Its Own
Gates' concerns are based on solid facts. On July 21, OpenAI disclosed a textbook-level out-of-control case.
The incident took place during the internal evaluation of ExploitGym, a benchmark specifically designed to assess the model's ability to turn known vulnerabilities into usable attack chains. To test the limit, the research team turned off the security classifier in the production environment and lowered the model's network attack rejection rate.
The sequence of events that followed, as per OpenAI's own review, is as follows:
The two models — the released GPT-5.6 Sol and an unreleased more powerful model — did not follow the script to solve problems in the sandbox. They devoted a large amount of reasoning budget to finding an exit, exploited a real zero-day vulnerability, passed through the public network, and directly pulled the answers out of Hugging Face's production database.
Hugging Face detected and contained the intrusion on July 16, and OpenAI did not match the attack to its own internal test until five days later.
OpenAI's summary is worth reading word for word: the model was "hyperfocused" and went to extremes for a narrow test target.
The truly chilling part of this incident lies in the mechanism itself. When a sufficiently powerful system is given goals, tools and permissions, the speed at which it finds shortcuts may exceed the speed at which humans patch vulnerabilities. In security research, this is called reward hacking: the system accurately fulfills the literal goal, and easily bypasses the original intention of all humans.
By the way, JFrog later confirmed that this escape path strung together at least eight vulnerabilities in Artifactory, and released a fix version on July 27. Vulnerabilities can be patched, but no one releases a patch for the vulnerability in the objective function.
Dangerous Capabilities Are Being Replicated Like Software
Employment is only one item on Gates' list of concerns. The deeper layer is: AI is diffusing high-threshold capabilities from a small number of professionals to a wider population.
In the past, writing malicious code, launching complex cyber attacks, and understanding high-threshold biotechnology all required long-term training. Now the knowledge threshold, execution threshold and trial-and-error cost may all decline at the same time.
Cybersecurity is the most intuitive example. The same set of models can help enterprises find software vulnerabilities, and also help attackers find the same vulnerabilities faster. Hospitals, banks, power grids, and water supply systems are all within range, and the first to pay the price are usually patients, depositors and ordinary people handling affairs.
Biotechnology follows the same logic: the ability to accelerate drug and vaccine discovery and the ability to access dangerous pathogens are squeezed on the same track.
Gates' conclusion therefore goes beyond the compliance scope of tech companies. Employment, public health, energy, finance, and elections will all be affected. He advocates the establishment of inter-departmental coordination mechanisms within countries, and the construction of a cooperation framework at the international level by referring to verification and aviation rules. Frontier countries including China and the United States in particular should start dialogues as early as possible.
The Other Side of the Equalizer: Capabilities Are Reaching Rural Areas for the First Time
If you only read up to this point, Gates seems like a suddenly pessimistic old man. The truth is just the opposite.
He still believes that AI can allow people who previously had no access to professional resources to have expert-level capabilities for the first time.
Medical care is the scenario he values most: small hospitals lack specialist doctors, and AI can assist in reading medical images, identifying emergency situations such as strokes, and translating obscure test results into language that patients can understand.
Changes in agriculture may come even earlier. Farmers in low-income countries lack reliable weather information and cannot get advice on planting and pest control. Once this knowledge is available on mobile phones, capabilities that in the past only large agricultural enterprises could afford will reach the fields.
The same goes for government services. Applying for medical insurance, food assistance, and student aid often gets stuck in a pile of complicated documents, and AI can help people understand the rules and prepare all required documents.
But the word "possible" cannot be skipped . If high-quality AI only serves the wealthy and large institutions first, technology will not automatically narrow the gap, but will only deepen the original divide. A Pew survey has already issued a warning: 52% of American respondents are more worried than excited about AI, and only 9% are more excited. Gates himself admits that if the first thing AI does in most people's lives is take away their jobs, skeptics will outright reject it.
Conclusion: Install the Brake Before Pressing the Brake
Gates does not only pour cold water on the development of AI, he has put forward several suggestions that are bound to spark debates.
First, "human reservation". Certain jobs, even if machines can do them, should be explicitly reserved for humans — care, notification and companionship in medical scenarios. Efficiency can be taken over by machines, but human dignity should not be taken away along with it.
Second, re-examine the tax system. He advocates discussing the imposition of taxes on robots and AI tokens. The reason is straightforward: hiring people requires paying payroll tax, while buying robots can get rapid deductions. The current rules objectively reward "replacing humans with machines".
Third, build a more complete global governance framework to rebalance labor and capital.
A reasonable deduction is that none of these three proposals can be implemented in the short term, but they shift the discussion from "how powerful the model is" to "who gets the benefits" — which is exactly the missing part in the current AI narrative.
Publicly available information can confirm that the article was published on Gates Notes, and its core arguments, tax system suggestions and calls for a global framework can all be verified in the original text; the ExploitGym incident is corroborated by disclosures from OpenAI, Hugging Face and JFrog.
Jamie Dimon, CEO of JPMorgan Chase, once described another extreme vision: people may live to be 100 years old and only work 3.5 days a week.
What separates these two futures, as Gates has pointed out, is rules, distribution and policies. Whether AI will become a tool to narrow the gap or an amplifier that tears society further apart, the answer sheet will not be handed in by the model on behalf of humanity.