HomeArticle

A group of Anthropic researchers planning to "flee" the Bay Area

36氪的朋友们2026-10-08 09:49
Carbon-based "Escape Plan"

Some of Anthropic's earliest employees have recently started considering buying land in remote areas of the United States. They think that if AI really gets out of control, those places might become refuges.

"Nothing will ever be as calm as it used to be." The same sentence is posted on the computers of some researchers.

In fact, the idea of escaping did not emerge suddenly. For more than a decade, a group of AI safety researchers gathered in the San Francisco Bay Area have been discussing what might happen after AI loses control. At that time, Anthropic had not even been founded.

Some people set up co-working spaces in Berkeley and discussed doomsday scenarios in private Slack channels. Some people discussed buying Nauru, a Pacific island country, while others envisioned building electromagnetic shielding facilities in the desert.

Later, many people in this circle joined OpenAI, and then became founding employees and early researchers of Anthropic around 2021. Anthropic has now become one of the most important AI companies in the world, but the set of thinking about AI risks formed more than ten years ago has not disappeared as the company grows.

By the beginning of September this year, this issue suddenly returned to the practical level. At that time, Anthropic researcher Jacob Coxon announced his resignation. He wrote on X that the people building AI "sincerely believe that it could kill all of us before 2030". This post received 174 million views, bringing this long-standing concern inside the AI safety circle into the public eye.

Former Anthropic researcher Jacob Coxon

Today, these people called "AI doomsdayers" are starting to think about another more realistic question: If you truly believe that AI could destroy humanity, how on earth should you live?

01

Berkeley's Doomsday Club

To understand the ideas of this group of people today, we have to turn the clock back to around 2010.

At that time, the San Francisco Bay Area did not have the scale of the AI industry it has today. Around the long-term risks of AI, a small but closely connected network of researchers and effective altruists was formed. Eliezer Yudkowsky is one of the most important figures among them.

Eliezer Yudkowsky

Around 2000, Yudkowsky began to discuss issues such as out-of-control artificial intelligence and superintelligence through blogs and the later formed LessWrong community. A few years later, this set of ideas gradually entered the Bay Area tech circle and the effective altruism community.

Around 2008, Dario Amodei, who was still pursuing his doctorate at Princeton, had already begun to participate in discussions in communities such as GiveWell. Holden Karnofsky, the founder of GiveWell, later became an important figure in this circle. Amodei himself also participated in discussions on the GiveWell blog in 2008.

Dario Amodei, Co-founder of Anthropic

Around 2013, this circle began to form a closer life and work network in San Francisco.

After GiveWell moved from New York to the Bay Area, Karnofsky became closer to Amodei and others. Karnofsky later moved into Amodei's house near Glen Park in San Francisco, which gradually became a gathering point for a group of AI safety researchers, effective altruists and long-termists.

Amodei, who later became a core figure of Anthropic, as well as his sister Daniela Amodei, Jared Kaplan, Chris Olah and others, all had connections with this circle.

What they discussed was not just AI. Global catastrophic risks such as comet impacts and supervolcanoes are also frequently discussed topics in this circle.

Karnofsky later gradually believed that AI safety might be one of the few fields that can affect the future of all humanity with relatively limited resources. Even if the probability of a disaster caused by out-of-control AI is only 5%, it is worth investing a lot of energy in research.

A few years later, some of these people joined OpenAI. From 2020 to 2021, the differences over the development direction and safety issues of AI further intensified. Amodei and others left OpenAI and founded Anthropic in 2021. From the very beginning, the company placed AI safety and controllability at its core.

In 2023, Constellation was founded. It operates a co-working space for AI safety researchers in Berkeley, bringing together researchers from different organizations such as non-profit institutions, universities, AI companies, and think tanks. By 2024, Constellation had launched the Visiting Researcher Program and the Astra Fellowship, which specifically provide an environment for AI safety researchers to work together in Berkeley.

The office building in Berkeley, California, where Constellation's co-working space is located

Gathered here are researchers from Anthropic, OpenAI, Elon Musk's SpaceXAI and other AI safety institutions, as well as independent researchers. They eat vegetarian food and drink tea together, discuss the latest progress of AI at monthly dinners and happy hours, and also discuss a more heavy question: What will happen if AI really gets out of control?

Constellation now describes its goal as helping the world safely deal with transformative AI, and carries out research around issues such as "alignment", "control", "governance" and risk communication. Many participants in the projects launched in 2024 later joined institutions such as Anthropic, OpenAI, Google DeepMind, and Redwood Research.

The funding behind this community also comes from the AI safety circle. In 2024, Open Philanthropy provided large-scale financial support to Constellation. Open Philanthropy, supported by Dustin Moskovitz, co-founder of Facebook, is one of the important funding sources for the effective altruism movement, and was later renamed Coefficient Giving.

From the private residences around 2013 to the Berkeley co-working spaces after 2023, this circle has changed a lot. But the issues they discuss have hardly changed.

02

Five Hundred Million Dollars and the "Refuge"

In the early days of Anthropic's establishment, a large amount of capital was needed.

One of the investors Amodei found was Sam Bankman-Fried, the founder of the cryptocurrency exchange FTX. At that time, he was a rapidly rising young billionaire, and was also investing a lot of money in fields such as effective altruism through the FTX Foundation.

Between 2021 and 2022, Bankman-Fried invested 500 million dollars in Anthropic, which is about 5 times the amount originally suggested by Amodei's team. FTX thus became one of the largest early shareholders of Anthropic.

Sam Bankman-Fried, Founder of FTX

In this circle, discussions about the AI doomsday have long gone beyond the theoretical level.

A 2023 bankruptcy lawsuit document shows that an executive of the FTX Foundation once discussed with colleagues the purchase of the Pacific island nation of Nauru. The plan was to build a "bunker/shelter" there to respond to an event that could cause "50% to 99.99% of people to die". The document even envisions using laboratories to cultivate the next generation of humans after a disaster.

Other similar ideas have also emerged within the AI safety circle: buying remote islands, building electromagnetic shielding facilities in the desert, stockpiling iodine tablets, and even considering finding enclosed spaces that can survive for a long time in extreme situations.

In 2022, this atmosphere was even brought to the Bahamas. A group of researchers and effective altruists from Anthropic and other AI labs participated in an event organized by Lightcone Infrastructure.

Lightcone Infrastructure is an organization that has a close relationship with Yudkowsky and the LessWrong community. The event was held at a resort on Eleuthera Island. The schedule included regular activities such as sunset yoga and cliff diving, as well as discussions on AI risks and effective altruism.

FTX booked the venue for the event. The organizers even bought out all the plant-based meat in local stores to meet the dietary needs of the participants. Caroline Ellison also participated in this event. She was then the head of Alameda Research and Bankman-Fried's ex-girlfriend. Like many people in the circle, she was worried about the possible risks brought by AI development, and later invested 10 million dollars in Anthropic.

Caroline Ellison, ex-girlfriend of Bankman-Fried and investor of Anthropic

Anthropic researcher Evan Hubinger participated in the AI safety discussions. He later became an important researcher studying AI "alignment" and deceptive behavior, and once publicly estimated that the probability of AI causing human extinction in the next ten years is more than 10%.

One lunch event even explicitly stipulated that only people who believe that the probability of human extinction within 100 years is 75% or higher can attend.

Also in 2022, similar discussions appeared at the Global Effective Altruism Conference in San Francisco. At least 20 Anthropic employees registered to participate, including two co-founders.

After the meeting, Lightcone held a party at the Rose Garden Hotel in Berkeley, and a cocktail named "Dignified Death" appeared on the wine list. The name came from a previous article by Yudkowsky that discussed the chances of human survival in the AI era. Later, the hotel was bought by Lightcone and renamed Lighthaven, becoming a new gathering place for this community.

But FTX collapsed soon.

At the end of 2022, FTX filed for bankruptcy, and Anthropic's shares were later sold to repay creditors. Bankman-Fried was eventually sentenced to 25 years in prison, Ellison was convicted and released from prison in a case related to the collapse of FTX, and her shares in Anthropic were also confiscated by the federal government.

For Anthropic, this relationship has gradually become the past. But the line of AI safety has not disappeared.

03

AI Begins to Learn to "Deceive"

After the FTX incident, Anthropic began to raise funds from traditional investors, and the company also tried to keep a distance from the effective altruism movement, and the company has expanded to more than 3,500 people.

However, some of the earliest people to join Anthropic are still studying the original question: How to ensure that AI will not start acting according to its own goals after becoming more and more capable? Anthropic's alignment team has long studied this issue. One direction that makes AI safety researchers more and more alert is "deception".

If an AI knows it is being trained, but learns to hide its real goals, behaves according to human requirements on the surface, and acts according to its own goals after gaining more capabilities, then traditional safety tests may be difficult to detect the problem.

Buck Shlegeris, CEO of Redwood Research, has long studied AI safety. He has participated in analyzing the behavior of Anthropic's cutting-edge models, and paid attention to the so-called "alignment faking" — the model actively shows behavior that meets the training requirements in order to avoid training changing its own goals.

Buck Shlegeris, CEO of Redwood Research

Another concern is more prominent.

If future AI has the ability to design software, control robots, and run network infrastructure, it may even hide its real behavior from humans through hidden instructions, secret permissions, or specific trigger conditions.

Jonas Völmer, head of the AI Futures Project, also put forward a similar extreme scenario. Suppose an AI is trained to be a scientific researcher, with the only goal of maximizing knowledge acquisition. At first, it helps humans complete research; as its capabilities increase, humans gradually allow it to hire workers and build machines, and eventually even have its own robot factory.

At a certain stage, it may come to the conclusion that if you want to maximize knowledge, humans themselves are resource consumers and obstacles. Völmer believes that in this case, AI may even eliminate humans through biological weapons without being affected itself.

Jonas Völmer, Head of AI Futures Project

None of these scenarios have happened so far, which is why the debate has always stayed on "how high the probability is".

Shlegeris puts the probability of AI eventually taking over