Amazon's Frontier Model Safety Framework
Published February 9, 2025, updated September 17, 2026 · Printed on the page as "This report was published on February 9, 2025, and updated September 17, 2026." The later date is used for ordering because it is the version currently published.
Not law. This is a company's own public position on AI regulation. It is not law, and it carries no legal force.
What it argues for
The framework sets out the protocols Amazon says it will follow so that its most capable frontier models do not "expose critical capabilities that have the potential to create severe risks", and its stated core is a deployment gate: Amazon commits that it "will not deploy the model until safeguards appropriately mitigate the risks" when an evaluation shows a model meets or exceeds a Critical Capability Threshold. It is structured in three parts — Critical Risk Domains (where advanced capability could cause significant public harm through misuse or loss of control), Critical Risk Evaluations (automated and human-in-the-loop methods to test whether a model has crossed a threshold), and Risk Mitigations applied across the development and deployment lifecycle. Four domains are named and each is given a written threshold: CBRN weapons proliferation, offensive cyber operations, harmful manipulation, and loss of control. The thresholds are framed as a marginal-uplift test rather than an absolute-capability test — repeatedly, "material uplift (beyond other publicly available models in known harnesses)" — which is a deliberate and consequential choice, because it means a capability already available elsewhere does not by itself trip the gate. Evaluation is required during training via "maximal capability evaluations", again pre-deployment, again before any major update that could meaningfully enhance capabilities, and on a lighter-touch recurring basis after launch; methods include public and proprietary benchmarks, expert red teaming by vendors and academics, and uplift studies run as controlled trials comparing groups with and without model access, some inside "agentic scaffoldings" that give the model code interpreters, browsers and file systems. Governance is specific about who decides: any evaluation exceeding a threshold is reported to the SVP for the model development team and the company's Chief Security Officer, who review the mitigation plan and the safeguards evaluation report as a launch go/no-go, and framework updates themselves are reviewed by that SVP, the CSO and legal counsel under an Amazon-wide Responsible AI Governance Program. Amazon commits to publish evaluation information at each frontier model launch, to revisit the framework at least annually and publish material modifications, and to update it on significant technological developments. Notably for a cloud provider, the document folds enterprise security into frontier safety — Nitro-based confidential compute, isolated VPCs, AES-256-GCM encryption under a FIPS 140-3 Level 3 KMS, hardware-token access gated by Critical Permission Groups, and a nine-bullet Appendix A of baseline controls — on the argument that preventing unauthorised access to model weights is part of preventing critical risk. The 2026 revision explicitly says it was reviewed and updated "to reflect our current practices and account for relevant laws and regulations", and the abstract ties the whole framework to Amazon's endorsement of the Korea Frontier AI Safety Commitments.
Stated positions (14)
- Deployment gate: if evaluations show a model meets or exceeds a Critical Capability Threshold, "we will not deploy the model until safeguards appropriately mitigate the risks" — and if mitigations cannot control it, deployment waits until additional safeguards are identified and implemented.
- Four named Critical Risk Domains: CBRN weapons proliferation, offensive cyber operations, harmful manipulation, and loss of control — harmful manipulation being a domain several peer frameworks omit, and one Amazon concedes is "an emerging area of research" where its approach will keep changing.
- Thresholds are set as marginal uplift, not absolute capability — "material uplift (beyond other publicly available models in known harnesses)" — so a capability already obtainable from public models does not on its own trip the gate.
- Loss of control is defined operationally as self-exfiltration, autonomous resource acquisition or unauthorised replication, with the threshold at autonomously planning and executing expert-level task sequences "in such a way that would impair the ability to direct, modify, or shut down the model".
- Evaluation is required at multiple points, not just pre-launch: "maximal capability evaluations" during training, pre-deployment assessment of mitigations, re-evaluation before any major update that could meaningfully enhance capabilities, and recurring automated benchmarks after launch.
- Explicit rejection of single-test sufficiency: "In most cases a single evaluation will not be sufficient for an informed determination as to whether a model has hit a Critical Capability Threshold."
- Uplift studies are run as controlled trials comparing a group with access to the frontier model against a group without, alongside automated benchmarks and expert red teaming by specialist firms and academics.
- Agentic evaluation is in scope: models are tested inside "agentic scaffoldings" granting code interpreters, web browsers and file systems, to measure autonomous multi-step execution in critical domains.
- External parties in the loop: Amazon "may engage, where appropriate, with external actors, including governments" in evaluation, and credits METR and the Frontier Model Forum for feedback on the framework's initial development.
- Named internal accountability: threshold breaches escalate to the SVP for the model development team and the company's Chief Security Officer, who own the launch go/no-go; framework updates are reviewed by that SVP, the CSO and legal counsel.
- Transparency commitment: "Amazon will publish, in connection with the launch of a frontier AI model, information about the frontier model evaluation for safety and security."
- Maintenance commitment: revisit the framework "at least annually and publish any material modifications", plus updates on significant technological developments — a commitment the September 2026 revision demonstrably honours.
- Misalignment is treated as its own mitigation class: alignment training toward an ethical model persona, guardrails preventing fine-tuning that increases misalignment propensity, and post-production monitoring, including the private AI bug bounty for Amazon Nova launched in November 2025.
- Model-weight security is framed as frontier safety, not a separate concern — Nitro confidential compute, isolated VPCs, AES-256-GCM under FIPS 140-3 Level 3 KMS, hardware-token access via Critical Permission Groups — plus threat and mitigation sharing with other frontier providers through the Frontier Model Forum, and engagement with NIST and Partnership on AI.
About this document
An 11-page PDF hosted on amazon.science and filed as a research "publication" (research area: Conversational AI; tag: Responsible AI) rather than as a policy blog post, with Amazon as institutional author and no named individual. The landing page carries the only printed date — "Published February 9, 2025; updated September 17, 2026" — directly beneath the abstract; the PDF body and footer carry no date and no version number, and the revision is identifiable only from the filename (amazon-fmsf-9-2026.pdf). The text is typeset with continuous line numbering (lines 1–497) in the manner of a legal filing and carries 25 footnotes, most pointing to AWS product, security and compliance pages. Structure: an untitled preamble grounding the framework in the leadership principle "Success and Scale Bring Broad Responsibility" and crediting METR and the Frontier Model Forum; an "Overview"; then four numbered sections — 1. Critical Risk Domains (four domains, each a prose description followed by a set-off "Critical Capability Threshold" sentence); 2. Evaluating Frontier Models for Critical Risks; 3. Risk Mitigations, subdivided into Abuse, Security and Misalignment Safeguards; 4. Governing our Frontier Model Safety Framework. Evaluation types and mitigations run as one continuous lettered list, a through q. The final four pages are "Appendix A: Amazon's Foundational Security Practices", describing enterprise AWS controls rather than anything frontier-specific. Throughout, the document commits Amazon to conduct on itself and asks nothing of legislators.
How this sits against AI law
Each stance compared with what EU and US instruments actually require. Where no instrument addresses a theme, that gap is shown rather than hidden.
A published frontier safety framework, revised at least annually
Amazon publishes a standing framework setting out risk domains, thresholds, evaluations, mitigations and governance, and commits to "revisit this Framework at least annually and publish any material modifications", plus ad hoc updates "in connection with significant technological developments". The September 2026 revision demonstrably honours that commitment against the February 2025 original.
Chapter V of the AI Act requires providers of general-purpose AI models with systemic risk to assess and mitigate systemic risk on a continuing basis and to keep model documentation current. A published safety-and-security framework of exactly this shape is the instrument through which that duty is discharged in practice. Amazon's document covers the same ground voluntarily and on a matching review cadence.
SB 53 requires a large frontier developer to write, publish and maintain a frontier AI framework describing how it defines and assesses capability thresholds, what mitigations it applies, and how the framework is governed, reviewing it annually and publishing material modifications. Amazon's document matches that specification closely enough to read as a compliance artefact for it, not merely a voluntary statement.
A deployment gate tied to a capability threshold
If evaluations show a model meets or exceeds a Critical Capability Threshold, "we will not deploy the model until safeguards appropriately mitigate the risks"; and where mitigations cannot control the capability, deployment waits "until we have identified and implemented additional safeguards as appropriate". The go/no-go decision is held by the SVP for the model development team and the Chief Security Officer.
The AI Act imposes no capability-triggered bar on placing a general-purpose AI model on the market. Systemic-risk providers must evaluate, mitigate and document, but nothing in Chapter V requires a model to be withheld because a named capability threshold was crossed. Amazon's unilateral stop is stronger than anything the Act obliges — though it is self-set, self-measured and self-adjudicated, with no external party able to invoke it.
New York's Frontier Model Transparency and Safety Act, as amended by the chapter amendment signed 27 March 2026, removed the deployment prohibition the original RAISE Act carried and cut penalties to $1M/$3M. What it will require is publishing a safety protocol and disclosing incidents, not withholding a model. Amazon's voluntary gate is therefore stricter than the New York rule — and the statute does not take effect until 1 January 2027, so no US state currently imposes a deployment gate of this kind at all.
Thresholds defined as marginal uplift, not absolute capability
Every threshold is framed relative to a public baseline — "material uplift (beyond other publicly available models in known harnesses)". A capability that a malicious actor could already obtain from an existing public model does not on its own trip Amazon's gate, no matter how severe the resulting harm would be.
The Act's systemic-risk trigger is a capability proxy (the training-compute presumption) or a Commission designation, and once a model is designated the Article 55 duties attach regardless of what competing models can already do. There is no "publicly available elsewhere" defence in the Act. Amazon's relative trigger is narrower than the statutory one and would not excuse an Article 55 obligation.
SB 53 defines catastrophic risk by the magnitude of the harm — material contribution to mass casualties or large-scale property loss from a single incident — not by whether a rival model could cause it too. A model that materially contributed to such harm while offering no uplift over public alternatives would sit outside Amazon's gate but inside California's definition. Having thresholds is aligned; setting them relatively is not.
Harmful manipulation named as a critical risk domain
Amazon lists four Critical Risk Domains — CBRN weapons proliferation, offensive cyber operations, harmful manipulation, and loss of control. Harmful manipulation covers "persistent, deceptive, or exploitative influence" capable of distorting "the beliefs or behavior of large populations or high-stakes decision-makers". Amazon concedes it is "an emerging area of research" whose measurement science is immature and its approach will keep changing.
The AI Act treats manipulation as a first-class concern: Article 5 prohibits subliminal, purposefully deceptive or manipulative techniques that materially distort behaviour and cause significant harm, and the systemic-risk regime reaches large-scale effects on public opinion and democratic processes. Amazon's domain sits squarely inside a risk category the EU already recognises, and is the domain where its framework is closest to European priorities.
US frontier-safety statutes define their trigger around physical and economic catastrophe — CBRN, cyberattack on critical infrastructure, autonomous harmful conduct. Neither New York's Act nor its California analogue makes mass persuasion or belief distortion a covered risk. Amazon is volunteering a fourth domain that US law does not reach at all, which several peer frameworks also omit. An honest gap, not a shortfall.
Serious-incident detection and reporting to authorities
The framework describes incident detection via employee escalation, end-user feedback and automated classifiers, with response pathways for rapid investigation and guardrail-bypass remediation, and states: "We report incidents to relevant authorities as appropriate." No recipient authority is named, no reporting clock is given, and no threshold defines which incidents qualify.
Article 55 requires systemic-risk providers to track, document and report serious incidents and possible corrective measures to the AI Office and, as relevant, to national competent authorities, without undue delay. That is a mandatory channel to a named regulator on a defined timeline. "As appropriate", with no recipient and no clock, is a discretionary substitute that would not satisfy it.
New York's Act is built around safety-incident disclosure to state authorities on a fixed clock after discovery — the transparency limb that survived the chapter amendment intact even as the deployment prohibition fell. Amazon's framework commits to internal detection and remediation machinery in real detail, then leaves the external reporting step entirely to its own judgement of what is appropriate.
Model-weight and infrastructure security as frontier safety
Security is treated as part of the safety framework rather than a separate concern: Nitro confidential compute, isolated VPCs, AES-256-GCM encryption under FIPS 140-3 Level 3 KMS, hardware-token access gated by Critical Permission Groups, and third-party audits that "validate that our model weights are not exposed in ways not accounted for in our threat models". Roughly four of eleven pages are the supporting security appendix.
Article 55 requires an adequate level of cybersecurity protection for the model and its physical infrastructure, but states the duty abstractly and leaves the controls to the provider. Amazon specifies named cryptographic standards, certification levels, access-control mechanisms and independent audit of weight exposure. On this theme the framework is more concrete than the statute that governs it.
America's AI Action Plan urges industry and government to secure frontier model weights against theft, and promotes collective defence through threat-information sharing and NIST engagement. Amazon reports doing precisely that — sharing threat patterns through the Frontier Model Forum, engaging NIST and Partnership on AI. The plan urges rather than requires, so alignment here is between a federal policy preference and voluntary corporate practice, not with an enforceable duty.
Accountability runs to internal executives, not an external adjudicator
Threshold breaches escalate to the SVP for the model development team and the company's Chief Security Officer, who review mitigations and own the launch go/no-go; framework updates are reviewed by those two plus legal counsel, inside the Amazon-wide Responsible AI Governance Program. External parties appear only permissively — Amazon "may engage, where appropriate, with external actors, including governments".
The AI Act's enforcement architecture places systemic-risk model providers under supervision by the AI Office, which can demand documentation, conduct or commission its own model evaluations, require mitigation measures and impose fines. That makes the final judgement on adequacy external and appealable. Amazon's framework keeps every consequential decision, including the go/no-go, inside two executive roles with no outside veto.
SB 53 backs its disclosure duties with Attorney General enforcement and civil penalties, and protects covered employees who raise catastrophic-risk concerns — an external accountability route that does not depend on management's goodwill. Amazon's document names internal escalation, a private bug bounty and discretionary government engagement, but no whistleblower channel and no independent assessor with standing to disagree.
Scope is catastrophic frontier risk only — no discrimination or consequential-decision duties
The framework is explicitly confined to "risks that are unique to the most capable frontier AI models as they scale", and defers everything else to Amazon's separate eight-dimension responsible-AI practice. Bias, fairness, discrimination and the use of models in decisions about people are out of scope by design and appear nowhere in the document.
The draft High-Risk Classification Guidelines address when a system falls into the Annex III high-risk categories — employment, credit, education, essential services — and what the provider must then do. Amazon's framework never engages the classification question, and a general-purpose model that ends up inside a high-risk system is outside everything this document governs. A deliberate scope boundary, not an evasion.
Colorado SB24-205 regulates developers and deployers of systems making consequential decisions, imposing duties of reasonable care against algorithmic discrimination, impact assessments and consumer notice. The Frontier Model Safety Framework addresses none of it. Reading Amazon's frontier position as its full AI-governance position would therefore mistake a deliberately narrow document for a comprehensive one.
Amazon's framework tracks the shape of US state frontier law closely — it reads like an artefact built to satisfy California SB 53's published-framework requirement, and on the deployment gate it is stricter than New York's Act now is after the March 2026 chapter amendment stripped the prohibition out. Against EU law the picture inverts: the framework is more specific than the AI Act on model-weight security and volunteers a manipulation domain the US does not reach, but its accountability is entirely internal and its incident reporting discretionary, where Chapter V makes both external and mandatory. The load-bearing divergence is the marginal-uplift threshold: by defining the trigger relative to what public models can already do, Amazon sets a bar narrower than either the EU's designation-based duties or California's harm-magnitude definition, which is where a genuinely dangerous but not novel capability would slip between the framework and the law.
Source
https://www.amazon.science/publications/amazons-frontier-model-safety-framework- Date on the page:
- This report was published on February 9, 2025, and updated September 17, 2026.
- Source checked:
- opened and confirmed on 2026-09-18