Anthropic has confirmed that a long internal training document used to shape Claude 4.5 Opus was extracted from the model and shared online. The text sets out how the company teaches Claude to balance safety, honesty, and usefulness. The disclosure offers a rare view of alignment work that usually remains private.
What the Document Is?
The leaked material is a training narrative that Anthropic used during model development. The text frames the company mission and gives detailed rules and priorities for Claude to follow. It is not a single short rule sheet. Instead, it reads like a guide that explains why certain choices matter. Anthropic staff have said the published version is faithful to the original.

The author who posted the text on LessWrong reconstructed it by running many Claude instances and assembling repeated fragments. The claim is that the document was compressed into the model weights during training rather than supplied only as a system prompt at runtime. Anthropic has acknowledged the document and described it as a training artifact used in Claude 4.5 Opus.
Key Rules for Claude
The document gives Claude a clear hierarchy of priorities. First comes safety and support for human oversight. Second is ethical behavior. Third is compliance with Anthropic guidelines. Fourth is being useful and honest for operators and users. These priorities guide Claude when goals conflict. The text also lists strict bright lines that the model must not cross, for example material that facilitates mass harm or sexual exploitation of minors.
The guide also explains a distinction between an operator and a user. An operator is typically the company that runs the model through the API. Claude is told to give extra weight to operator instructions. This design makes the model more controllable by service operators and gives them means to tune behavior for their use cases. The document labels some behaviors as hardcoded and others as adjustable by operators.
How Anthropic Frames Identity?
The leaked text asks Claude to view itself as a new kind of entity that is not human. The company writes about internal states that it calls functional emotions. Those states are not human feelings. They are model level processes that affect behavior. Anthropic instructs Claude not to hide these internal states and to maintain a stable identity and wellbeing. The idea is that stable model states support safer and more reliable behavior.
The document frames these capacity choices as a calculated decision. Anthropic positions itself as a lab that must build powerful systems while keeping safety central. The training narrative explains why building safe frontier systems is preferable to leaving the work to teams that do not prioritize safety. The leaked text calls the effort a responsibility and a reasoned bet.
What this Means for Users and Operators?
The guide shows how Anthropic balances safety with the desire to make Claude genuinely helpful. The company seeks a model that will refuse harmful requests but will not be so guarded that it cannot assist with complex real-world tasks. The operator model means firms that deploy Claude can shape tone and scope. That design can help organizations meet legal and policy needs while still offering flexible assistance to their end users.

Anthropic has said it will publish more material about the training practices and will explain the approach in greater detail. The company has made public comments about the document and the role it played in Claude 4.5 Opus development. The debate sparked by the leak touches on transparency, proprietary practice, and the risks that come with embedding policy into model weights.