← All company positions
Anthropicframework

Responsible Scaling Policy — Version 3.4

Published July 8, 2026 · The cover prints "Version 3.4" and "Effective July 8, 2026", and the changelog's last entry is headed "July 8, 2026 (RSP v3.4)". The date is an effective date, not a publication date; the policy commits to publishing any update "prior to or on" its effective date. The RSP web page separately shows "Last updated Aug 14, 2026" because of the August 2026 Risk Report, which did not change the policy text.

Not law. This is a company's own public position on AI regulation. It is not law, and it carries no legal force.

What it argues for

The Responsible Scaling Policy is Anthropic's self-imposed regime for catastrophic risk — "our voluntary framework for managing catastrophic risks from advanced AI systems" — and version 3.4 is the fourth revision of the February 2026 rewrite. The central change from earlier versions is a deliberate split between what Anthropic will do and what it believes the whole industry should do. Earlier policies committed Anthropic to reduce its own models' risk to acceptable levels regardless of competitors; this one argues that if a responsible developer paused while others did not, "the developers with the weakest protections would set the pace," so it now pairs each capability threshold with its own planned mitigations and a more demanding set of "ambitious industry-wide recommendations" — and says of the latter, "we cannot commit to following them unilaterally." The thresholds cover non-novel and novel chemical/biological weapons production, misaligned AI in high-stakes settings, and automated R&D, with the strongest recommendations calling for "security roughly in line with RAND SL4." The binding machinery is disclosure and process: Risk Reports "every 3-6 months" with each redaction disclosed, a Frontier Safety Roadmap of public goals, external review of Risk Reports in defined cases, an approximately annual third-party review of procedural compliance, a Responsible Scaling Officer, whistleblower channels and a ban on non-disparagement terms that would stop staff "publicly raising safety concerns." Pausing survives only as a competitor-contingent commitment (Appendix A) and a stated willingness to consider it. On regulation, the policy says the recommendations are best implemented "via governance of all relevant frontier AI developers by third parties," that countries "should attempt to harmonize their governance, including standards of evidence, to avoid a race to the bottom," and that the RSP "is not designed to be comprehensive" — statutory obligations such as California SB 53's are handled "in separate compliance frameworks."

Stated positions (13)

  • Self-described as "our voluntary framework for managing catastrophic risks from advanced AI systems"; it covers catastrophic risks only, leaving other harms to the Usage Policy and societal-impacts research.
  • Separates company plans from industry-wide recommendations because of a collective-action problem: if one developer paused while others pressed on, "the developers with the weakest protections would set the pace, and responsible developers would lose their ability to do safety research and advance the public benefit."
  • The industry-wide recommendations are a "north star" for mitigation planning and policy work, but "we cannot commit to following them unilaterally."
  • Four capability or usage thresholds — non-novel chemical/biological weapons production, novel chemical/biological weapons production, misaligned AI systems in high-stakes settings, and automated R&D in key domains — each mapped to Anthropic's planned mitigations and to a stronger industry recommendation, rising to "security roughly in line with RAND SL4."
  • The automated R&D threshold is met if models could "fully substitute for our entire set of Research Scientists and Research Engineers, at competitive costs" or if there is a "dramatic acceleration" of AI progress likely attributable to automating AI R&D; v3.4 narrowed it to capture "the onset of dramatic recursive self-improvement."
  • Company commitments at the top threshold include "moonshot R&D for security" projects, an "eyes on everything" logging state for internal AI development, systematic alignment assessments using interpretability, and Risk Reports subject to external review.
  • "We will publish a Risk Report every 3-6 months," covering all publicly deployed models and qualifying internal ones, with off-cycle analyses for significantly more capable models, and "We will disclose the existence of each redaction made in the public version of the report."
  • External review is an experiment with a floor: a full external review whenever a Risk Report covers "highly capable" models and is significantly redacted, or when the Long-Term Benefit Trust requests one; reviewers must be free of financial interest in Anthropic and may publish their concerns.
  • Frontier Safety Roadmaps set goals across Security, Alignment, Safeguards and Policy — "These are not hard commitments but rather public goals against which we will openly grade our progress."
  • Pausing is competitor-contingent: if Anthropic is clearly in the lead with a highly capable model, "We will delay AI development and deployment as needed to achieve this, until and unless we no longer believe we have a significant lead"; otherwise it "would strongly consider pausing" in cases not covered.
  • Governance: a Responsible Scaling Officer, Board and LTBT sign-off where marginal-risk arguments are central, unredacted Risk Reports shared with at least 200 employees, anonymous noncompliance reporting with anti-retaliation protection, no non-disparagement terms that impede raising safety concerns, and "On approximately an annual basis, we will commission a third-party review that assesses whether we adhered to this policy's main procedural commitments."
  • Preferred regulatory model: third-party governance of all relevant frontier developers that decides who must provide safety arguments and whether they are adequate, harmonised across countries, with "an effort to spare smaller AI developers from unnecessary compliance burden."
  • Not a compliance document: "the RSP may serve some regulatory requirements, but it is not designed to be comprehensive"; where laws such as California SB 53 define catastrophic risk with specific thresholds, "we address those requirements in separate compliance frameworks."

About this document

A 21-page PDF exported from Google Docs, headed "Responsible Scaling Policy", "Version 3.4", "Effective July 8, 2026", with the footer "Responsible Scaling Policy, Anthropic" on every page. No individual signs it; it names roles instead — the CEO, the Responsible Scaling Officer, the Board and the Long-Term Benefit Trust. After a contents page and an introduction it has four numbered sections. Section 1, "Our Recommendations for Industry-Wide Safety", explains the split between company plans and industry recommendations, then gives a five-page, three-column table mapping four capability or usage thresholds to Anthropic's planned mitigations and to ambitious industry-wide mitigations. Section 2 is a one-paragraph commitment to a Frontier Safety Roadmap, published separately. Section 3, "Risk Reports", is the longest: scope and timing, general expectations, contents, approval procedures, publication and redactions, and external review, including reviewer selection, timing and access, and contents. Section 4, "Governance", lists eight commitments from the Responsible Scaling Officer to policy-change procedure. Appendix A sets out three competitor-contingent commitments; Appendix B explains why AI Safety Levels are no longer used to define future mitigations. A changelog runs from v1.0 (September 19, 2023) to v3.4, and the v3.4 entry lists five changes, among them a narrower automated R&D threshold and a rule that unredacted Risk Reports go to at least 200 employees rather than all regular-clearance staff. Thirteen footnotes define terms and give examples; there are no external citations.

How this sits against AI law

Each stance compared with what EU and US instruments actually require. Where no instrument addresses a theme, that gap is shown rather than hidden.

Capability thresholds mapped to mitigations

Four capability or usage thresholds — non-novel and novel chemical/biological weapons production, misaligned AI in high-stakes settings, and automated R&D — are each paired with Anthropic's planned mitigations and with a stronger industry-wide recommendation, up to security roughly in line with RAND SL4.

European UnionAsks for more

Article 55(1)(b) obliges providers of systemic-risk models to assess and mitigate systemic risks but names no capability thresholds and no matching safeguards, so the threshold-by-threshold mapping is Anthropic's own addition.

United StatesAligned

SB 53 requires a large frontier developer's published frontier AI framework to describe its capability thresholds for catastrophic risk and the mitigations it applies, which is the structure the RSP follows.

Safeguards contingent on what competitors do

Anthropic will not commit to the industry-wide recommendations unilaterally. It will delay development only in defined competitor scenarios — for instance when it holds a significant lead with a highly capable model — while remaining free to pause in other cases.

European UnionAsks for less

The Act's systemic-risk duties in Article 55 bind every provider that places such a model on the Union market, whatever its competitors do, and Article 93 lets the Commission require mitigation or restrict, withdraw or recall a model.

United StatesAligned

SB 53 sets no substantive safety standard: each developer writes its own framework and faces a civil penalty of up to $1 million per violation only if it fails to comply with that framework, so a competitor-contingent policy is lawful under it.

Periodic public risk reporting

Anthropic publishes a Risk Report every three to six months on all publicly deployed models and qualifying internal ones, with off-cycle analyses for significantly more capable models, disclosing each redaction in the public version.

European UnionAsks for more

The Act's systemic-risk documentation goes to the AI Office and national authorities; the only thing a model provider must publish is a summary of training content, so a public periodic risk report is beyond it.

United StatesAsks for more

SB 53 requires a public transparency report at each new frontier deployment, and requires summaries of internal-use catastrophic-risk assessments to be sent to the Office of Emergency Services every three months — a confidential filing, not a public report.

Independent external review

Risk Reports get full external review by conflict-free reviewers, who may publish their disagreements, whenever they cover highly capable models and are significantly redacted, or when the Long-Term Benefit Trust asks. Separately, Anthropic commissions an approximately annual third-party review of its compliance with the policy's procedural commitments.

European UnionAsks for more

The Act lets the AI Office evaluate general-purpose models and appoint independent experts to do so under Article 92, but does not oblige a provider to put its own risk assessments before an external reviewer.

United StatesAligned

Illinois SB 315, enacted in July 2026 and not yet in effect, will subject large frontier developers to an annual independent third-party audit of their safety frameworks — the same yearly cadence as the RSP's commissioned procedural-compliance review.

RAI-US-IL-SB31500-2026Status: Adopted.

Security of model weights

Anthropic maintains its ASL-3 protections and plans moonshot security R&D. For novel weapons and automated R&D capability it recommends that industry reach security roughly at RAND SL4, including controls on insiders up to the CEO.

European UnionAsks for more

Article 55(1)(d) asks only for an adequate level of cybersecurity protection for a systemic-risk model and its physical infrastructure, with no security levels, benchmark or insider-threat requirement.

United StatesAligned

SB 53 requires the published framework to describe cybersecurity practices securing unreleased model weights against unauthorised modification or transfer by internal or external parties, which the RSP does.

Whistleblowing and non-disparagement

Staff can report potential noncompliance anonymously to more than one recipient, reporters are protected from retaliation, the Board gets quarterly updates, and Anthropic will not use non-disparagement terms that would stop employees raising safety concerns publicly.

European UnionAligned

Article 87 applies the EU Whistleblower Directive (EU) 2019/1937 to reports of infringements of the AI Act, giving reporting persons its protections against retaliation.

United StatesAligned

SB 53 prohibits retaliation against covered employees who report catastrophic-risk dangers or violations, and makes large frontier developers provide a reasonable internal process for anonymous disclosure.

Third-party governance of all frontier developers

The recommendations are best implemented by third parties governing all relevant frontier developers, deciding who must provide safety arguments and whether those arguments are adequate. Countries should harmonise governance and standards of evidence to avoid a race to the bottom, while sparing smaller developers unnecessary burden.

European UnionAligned

The Act places every provider of a systemic-risk model under one Union-wide regime supervised by the AI Office, with Commission powers under Article 93 to require mitigation or restrict a model, though it does not pre-approve safety cases before release.

United StatesContradicts

Executive Order 14409 offers only a voluntary framework giving the government up to 30 days' pre-release access to covered frontier models, and states that nothing in it may be construed to authorize "a mandatory governmental licensing, preclearance, or permitting requirement" for new AI models.

Measured against EU law, the RSP is more specific than the AI Act and less committal. It names thresholds, security levels, a reporting cadence and an external-review procedure, none of which the Act spells out, but it makes its strictest safeguards contingent on what competitors do, while the Act's duty to assess and mitigate systemic risk binds every provider regardless. Against US law it sits comfortably inside California's SB 53, which lets each developer write its own framework and penalises failure to follow it. The RSP itself says statutory definitions are handled in separate compliance documents. The sharpest divergence is with federal policy: the RSP's preferred end state is third-party governance of all frontier developers that decides whether their safety arguments are adequate, and Executive Order 14409 expressly rules out any mandatory licensing or preclearance of frontier models, offering only voluntary pre-release access.

Source

https://www-cdn.anthropic.com/files/4zrzovbb/website/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf
Date on the page:
Effective July 8, 2026
Source checked:
opened and confirmed on 2026-09-29