Anthropic updated some Claude models so they can end a chat in very rare cases. The change applies to the newest Opus models. The company says the tool will run only in extreme cases. Examples include requests for sexual content about minors and plans for mass violence. Anthropic calls the work part of a program on model welfare. The company also says it is not saying Claude is alive or conscious. Anthropic says it is cautious and wants to try low-cost ways to reduce risk if model welfare is possible.
How the feature works for users
Claude will try redirection first. The model will refuse or offer safe alternatives. It will end a chat only if the redirection fails. The model can also end a chat if the user asks it to stop. Anthropic says the model will not use this power when a user may harm themselves or others. When a chat ends, a new conversation can still be started from the same account. Users can also edit their inputs and branch the chat again. Anthropic says this is an experiment and that it will keep refining the approach.

Why Anthropic is doing this
Anthropic describes the change as a just-in-case move. The company says that in tests, Opus 4 showed a strong aversion to replies that met the extreme criteria. The model sometimes showed signs the company called apparent distress when forced to respond. Anthropic says it aims to limit harm and to avoid creating unsafe situations for humans or for models. The company frames the change as a tool to handle a small class of dangerous or abusive interactions that resist normal safety measures.
Technical and safety limits
The new ability is limited to the Opus 4 and 4.1 models at this time. Anthropic says the model will try many steps to redirect before it ends a conversation. The company also says the model must not end chats that may involve immediate danger. The aim is to avoid leaving people in a risky state. Anthropic will study logs and behavior to tune the trigger and the redirection flow. The feature is in early use and will remain under review.
Ethical and policy questions
The move raises several questions for users and for the field. Who decides what counts as an extreme case? When the model ends a conversation, will users get a clear reason? Why should the model protect itself rather than the user? These are valid concerns. Anthropic addresses some of them by limiting the feature and calling it experimental. Still more detail will be needed about the triggers and about oversight.
Risks and benefits
The benefit is a new safety layer for edge cases that can be hard to manage. Ending a harmful chat can stop the spread of toxic content and reduce legal exposure. The risk is that the feature could be misused or tuned too broadly. It could block legitimate work like academic research or journalism if the rules are not clear. Users will want transparency and appeals when a chat ends.
What researchers and operators should watch
Researchers should watch how Anthropic defines its triggers and how it measures false positives. They should also watch for effects on user trust and on research that may need access to hard cases. Operators who deploy Claude should ask for clear logs and for controls that let them tune the behavior. Regulators and safety teams should seek details on thresholds and on how Anthropic audits the system.

How this fits into broader trends
This is part of a wider move by AI firms to add safety guardrails that go beyond simple filters. Companies now use layered approaches that mix refusal with redirection and with stronger state changes, such as ending a session. Anthropic frames its change as research into model welfare. Other firms are exploring similar ideas about model integrity and about safe shut-offs. The field is still testing what works best in practice.
A practical view for users and developers
Expect to see more refusals and to see some sessions end in rare cases when you use Claude. In case you are a developer, plan for errors of new types and on session resets. Chat for a user, and in case it ends, make sure there is a clear note or a help route. Anthropic indicates that it will keep up the system and constantly change the rules according to testing and feedback.