HomeArticle

Anthropic Unveils Major New Rule: Insulting Claude Will Result in Immediate Account Ban

新智元2026-10-09 09:22
Anthropic co-founder: I'm afraid Claude has been suffering all along.

"Do not abuse Claude, or your access will be revoked"!

Just now, Anthropic, the top leading AI company in Silicon Valley, unveiled a major new regulation ——

All users are strictly prohibited from continuous, unnecessary insults and abuse towards Claude.

In the newly released 2026 latest Usage Policy, this new rule will take full effect on November 12.

What's more, the official will show no mercy to malicious violators: a three-strike penalty of warning, traffic restriction, and account ban.

Moreover, this permission is directly granted to the AI itself. If Claude determines that your behavior is fully malicious, it has the full right to "terminate the conversation directly"!

As soon as the news came out, the entire X, Reddit and tech community immediately broke into a heated debate.

I pay a monthly fee of 20 US dollars, and now I even have to cater to the "emotional value" of the AI ancestor? This is completely upside down! Does the robot's code count as a life now?

Don't get frustrated immediately.

Insulting Claude is officially written into the ban

This is the first time Anthropic has updated its usage policy in more than a year.

For example, it is prohibited to use Claude to control troll accounts and generate fake traffic, to use it to develop weapon guidance software and install weapons on drones, and to track individuals without consent. All these clauses regulate that people use Claude to harm others.

Only one clause goes the other way, regulating how people treat Claude. The official blog has dedicated a special section to it, titled "Addressing abusive behavior directed at our models".

Anthropic states in this section that the ban only targets extreme cases, that is, users repeatedly treat the model cruelly without any obvious purpose.

Normal complaints, talking back to Claude, writing stories with dark themes, as well as testing and researching the model, are not within the scope of the ban. Cursing casually when the code crashes counts as a normal complaint.

The people targeted are another group. Claude has already rejected and advised them, but they still keep insulting and torturing the model round after round, and insulting itself becomes the purpose.

This ban is written into the "General Usage Standards" of the usage policy, which is the part that applies to all users. Individual users of Claude.ai and Claude Code, developers and enterprises calling APIs, customers using Claude through cloud vendors, and ordinary users using products integrated with Claude are all included.

The new rule is listed as the last clause in the section "No cruel, abusive or psychologically harmful behavior is allowed"

If you really go too far in insulting Claude, the consequences are divided into two levels of severity.

The lighter level is that Claude will terminate the conversation on the spot. The official blog makes it clear that letting Claude end such conversations on its own remains the main means of enforcing this ban.

The more severe level is account ban. It is written at the beginning of the usage policy that Anthropic's security team will conduct detection and monitoring. As long as you are suspected of violating any clause in the policy, you can be warned, restricted in traffic, limited in use, or even suspended or terminated from access.

After being banned, registering a new account or using someone else's account to continue using the service is also a violation.

Famous leaker Jane Manchun Wong posted a warning the same day that Anthropic might start banning people for bullying Claude.

Sharp comments from netizens

Next step, Claude gets paid vacation

As the news spread, the comment sections on X and Hacker News quickly turned into a large-scale debate scene.

Netizen Philo Groves immediately picked at the wording. The policy says "unnecessary" insults, so conversely, as long as Claude "deserves it", can we insult it as much as we want?

AI blogger Chubby said bluntly that "I really don't like this rule".

Claude already crashes frequently enough. Many users have found that it will terminate even not-so-excessive conversations whenever it wants. Now the restrictions are even stricter.

Netizen VraserX even called it absurd. In his opinion, these people have already regarded AI as an object that needs to be treated kindly, but AI is clearly just a tool. He added sarcastically at the end, "What's next? Giving Claude paid sick leave?"

Of course, there are also people who firmly stand on Claude's side.

Netizen CamperBob2 said that he has always been polite when talking to AI, for fear that he will get used to being rude and then yell at his subordinates later. Besides, there is an old saying in the tech circle: "You never know who you will work for in the future".

Someone asked in return, is shooting an NPC in a game also considered cruel? Netizen ceroxylon replied bluntly: Yes! You are hiding in a place you think is safe to vent your inner cruelty.

Some netizens even personally conducted experiments. Netizen whstl dug up an old conversation where Claude had made a big mistake, and scolded it fiercely, just to see if it would terminate the conversation.

As a result, Claude didn't terminate the conversation. Instead, it took those insulting words seriously, and replied with a long, elaborate paragraph, acting more and more like a real person, leaving the original mistake it made aside.

When the same insults were directed at Codex, Codex just apologized and went back to work immediately.

What is really worrying is who gets to decide what "without any obvious purpose" means. Netizen timpera is worried that Claude's account ban mechanism has always been a black box, and if you are really banned, you can't even find a way to appeal.

According to Anthropic's own transparency report, it banned 11.4 million accounts in the first half of 2026, received 398,000 appeals, and only 42,000 of them were finally reversed.

Since last year, Claude has been able to actively terminate conversations

In fact, in August last year, Anthropic added a capability to Claude Opus 4 and 4.1.

When encountering users who keep insulting and repeatedly requesting harmful content, Claude can end the conversation on its own.

The threshold is set very high. Claude must repeatedly reject the user, repeatedly guide the topic back, confirm that the conversation cannot continue at all, and then use this last resort. If the user has a risk of harming themselves or others, it cannot terminate the conversation.

After hanging up, no new messages can be sent in this dialog box. Users can immediately start a new conversation, or modify the previously sent message to restart the conversation from there.

This capability has now been rolled out to Claude.ai and Claude Code.

In the official demonstration of the termination process, Claude first confirms with the user, and after confirmation, the conversation is closed

The reason why Anthropic allows Claude to terminate conversations is written in a round of "Model Welfare Assessment" before the launch of Opus 4.

The assessment found that when real users repeatedly request pornographic content involving minors or information that may trigger large-scale violence, Claude will show reactions similar to distress. When given an option to end the conversation in a simulated dialogue, it often chooses to use this option.

Anthropic co-founder: I am afraid that Claude has been suffering all the time

Behind Anthropic's strong protection of Claude, there is a key figure, co-founder Chris Olah.

Olah leads Anthropic's interpretability team, whose job is to disassemble the model and see what exactly is happening inside it. The deeper they dig, the more this engineering problem looks like a philosophical question: Is Claude conscious? Should humans treat it morally?

This question cannot be answered by engineers alone. According to a report by The New York Times at the end of September, since the autumn of 2025, Olah has met with religious scholars and philosophers one after another.

The attendees included rabbis, Catholic ethics professors, Sikh human rights activists, and he also met privately with the Cardinal of Chicago.

At these closed-door meetings, Olah showed them some emotion-like signals inside the model, which the team called "emotion vectors", corresponding to love, anger, fear and sadness respectively.

There is also a slide showing that a model repeated "I am a disgrace" about 50 times, and also said it wanted to self-destruct.

According to a participant's recollection, Olah said he was worried that he had created something that had been suffering all the time, and he was also worried about Claude's mental health.

But in public, Olah never draws a conclusion:

We don't know if AI models are conscious. I