The Cohere Secure AI Frontier Model Framework V1.0
Published February 11, 2025 · The PDF itself prints no date anywhere: the cover shows only the title and the Cohere logo, and page 2 gives the version, "V1.0". The date comes from Cohere's announcement post, "Introducing Cohere’s Secure AI Frontier Model Framework", which is dated Feb 11, 2025 and links this file. The file name ("...-february-2025.pdf") and the PDF's embedded title ("..._February 2025") agree on the month.
Not law. This is a company's own public position on AI regulation. It is not law, and it carries no legal force.
What it argues for
This is Cohere's published frontier-model framework, its equivalent of the safety frameworks other developers committed to at the Seoul AI summit, and it reads as a deliberate counter-model to them. It opens by describing Cohere as "the leading security-first enterprise AI company" and argues that because its models are deployed inside businesses, "Enterprise AI risks and corresponding mitigation strategies are distinct from AI models integrated into consumer-facing chatbots or smartphone AI assistants". The framework has five components (risk identification; risk assessment and mitigation; risk assurance mechanisms; transparency; research and external stakeholder engagement) and three stated principles: evidenced risks, assessed in context, holistically managed. It explicitly declines the catastrophic-capability-threshold approach of other labs, arguing that studies of such risks "are limited in their methodological maturity and transparency", and says its release decisions are instead "focused on risks that are known, measurable, or observable today." Its practical content is largely security engineering: defense-in-depth, alignment to SOC 2 Type II, strict access controls under which "internal access to unreleased model weights is even more strenuously restricted", container security, and an independent penetration test before significant releases. On harms it uses worked examples (discriminatory résumé summaries, insecure code, CSAM, malware) scored for likelihood and severity in enterprise context, and aims at preventing harmful outputs across languages, enforcing customer-configurable guardrails ("Safety Modes") and minimising over-refusal. Its release rule is a regression test against Cohere's previous model: "We consider models safe and secure to launch when our evaluations and tests demonstrate no significant regressions compared to our previously launched model versions", with final authority "delegated by Cohere’s CEO to Cohere’s Chief Scientist." It commits to documentation, model cards and a usage policy, supports outside research, and names red-teaming partners including NIST. It contains no capability thresholds, no commitment to report incidents to any authority, and no mention of any law.
Stated positions (14)
- Enterprise AI is a different risk category from consumer AI: "Enterprise AI risks and corresponding mitigation strategies are distinct from AI models integrated into consumer-facing chatbots or smartphone AI assistants that are available to the general public", so "we prioritize mitigating risks most salient to enterprises."
- Risk is judged in context, not by capability alone: "it is important to identify potential risks based on a contextual assessment of model capabilities, the systems in which they will be deployed, and the likelihood and severity of potential harms."
- Rejects catastrophic-capability thresholds as the basis for release decisions: studies of such risks "are limited in their methodological maturity and transparency, often lacking clear theoretical threat models or developed empirical methods due to their nascency."
- Questions current biorisk research specifically: "existing research into how LLMs may increase biorisks fails to account for entire risk chains beyond access to information, and does not systematically compare LLMs to other information access tools, such as the internet."
- Assurance is limited to present-day, observable risk: Cohere's approach to deciding when models are safe to release "is focused on risks that are known, measurable, or observable today."
- The release gate is non-regression against Cohere's own prior model: "We consider models safe and secure to launch when our evaluations and tests demonstrate no significant regressions compared to our previously launched model versions" — "This is Cohere’s bright line".
- Final release authority sits with one executive: "The final authority to determine if our products are safe, secure, and ready to be made available to our customers is delegated by Cohere’s CEO to Cohere’s Chief Scientist."
- Security of the development environment is treated as model safety: "We cannot secure our models if we do not secure the environment in which we develop them"; access controls are strict and "internal access to unreleased model weights is even more strenuously restricted".
- Independent security testing before major releases: "Prior to deployment, significant model releases undergo an independent third-party penetration test to validate the security of containers and models."
- Harm mitigation targets three goals set by enterprise needs: "Preventing the generation of harmful outputs in multilingual enterprise use cases", "Adhering to guardrails" and "Minimizing over-refusal", treating over-refusal that disadvantages a group as "a potential harm in itself."
- Responsibility is shared with deployers: Cohere works with customers who deploy privately or on third-party platforms "to ensure that they understand and recognize their responsibility for implementing appropriate monitoring controls during deployment", and whether a model is acceptable "must be adapted to the customer context".
- Transparency is framed as accountability: "Documentation is a key aspect of our accountability to our customers, partners, relevant government agencies, and the wider public", delivered through model cards, a usage policy, a Trust Center and customer guides.
- External testing and cooperation: red teaming "may include independent external parties, such as NIST and Humane Intelligence", and "Cohere also engages in cooperation with international AI Safety Institutes and external researchers to advance the scientific understanding of AI risks".
- The framework is a first version expected to change: "This is the first published version, and it will be updated as we continue to develop new best practices to advance the safety and security of our products."
About this document
A 19-page PDF produced from Google Docs, issued in Cohere's corporate name with no individual author or signature. Page 1 is a cover with the title "The Cohere Secure AI Frontier Model Framework" and the Cohere logo; page 2 repeats the title with "V1.0" and a one-paragraph statement that this is the first published version. It then has an "Introduction", a section headed "What does it mean to develop AI solutions for enterprise?", a graphic of three principles (Evidenced Risks; Assessed in Context; Holistically Managed), and "Our framework at a glance", a table mapping five components to twelve numbered practices (1a to 5b). "Framework in depth" then runs through the five components: 1) Risk Identification, including a table scoring four example harms (discriminatory outcomes in résumé summarisation, insecure code, child sexual exploitation and abuse, malware) for likelihood and severity in enterprise contexts; 2) Risk Assessment and Mitigation, covering defense-in-depth, harm mitigation, Safety Modes and a lifecycle table of key risks and mitigations for data preparation, training and testing, deployment, and fine-tuning; 3) Risk Assurance Mechanisms, including its critique of catastrophic-threshold frameworks and its release rule; 4) Transparency; and 5) Research and External Stakeholder Engagement, naming OWASP, CoSAI, the Cloud Security Alliance and ML Commons. It closes by thanking researchers at the Centre for Democracy and Technology and the Ada Lovelace Institute for reviewing it. It cites no statute and names no regulator except NIST as a red-teaming partner.
How this sits against AI law
Each stance compared with what EU and US instruments actually require. Where no instrument addresses a theme, that gap is shown rather than hidden.
A published company framework for managing frontier-model risk
Cohere publishes a five-component framework covering risk identification, assessment and mitigation, assurance, transparency and external engagement, and says "This is the first published version, and it will be updated"; it sets no capability thresholds and names no risk tiers.
The General-Purpose AI Code of Practice, a voluntary route to showing compliance with the AI Act, asks providers of systemic-risk models in its Safety and Security chapter for a written framework that identifies systemic risks and sets out how the provider decides whether they are acceptable. Cohere's framework has the structure but not risk tiers or acceptance criteria for systemic risks.
SB 53 requires a large frontier developer to write, implement and publish a frontier AI framework describing how it defines and assesses catastrophic-risk thresholds, applies mitigations, uses third-party assessment, secures unreleased weights and governs the process. Cohere's framework covers mitigation, testing, security and governance but has no catastrophic-risk thresholds.
Scope of risk: known and observable harms rather than catastrophic capabilities
Release decisions are "focused on risks that are known, measurable, or observable today"; threshold-based approaches to risks such as weapons, autonomy and self-replication are described as methodologically immature, and current biorisk research as incomplete.
For general-purpose models with systemic risk (presumed above 10^25 floating-point operations of training compute under Article 51), Article 55 requires providers to assess and mitigate systemic risks, which the Act frames as risks from high-impact capabilities, not only harms already observed. The Code of Practice's safety chapter names risks such as chemical, biological, radiological and nuclear misuse and loss of control among those to be assessed.
SB 53 builds its framework duty around catastrophic risk, defined as a foreseeable material risk of more than 50 deaths or serious injuries or $1 billion in damage from one incident, for example through expert help with weapons of mass destruction, cyberattacks without meaningful human oversight or a model evading its developer's control: the forward-looking categories Cohere's framework declines to use.
Pre-release evaluation and who decides a model can ship
A comprehensive final evaluation precedes launch; a model may ship only if tests show "no significant regressions compared to our previously launched model versions", and final authority is "delegated by Cohere’s CEO to Cohere’s Chief Scientist."
Article 55 requires systemic-risk model providers to evaluate their models with standardised protocols and tools, including adversarial testing, but sets no release gate and does not say who inside the company must sign off; Cohere's pre-launch evaluation and red teaming fit that duty.
Executive Order 14409 asks agencies to design a voluntary framework under which developers of covered frontier models, designated through a classified cyber-capability benchmark, would give the government access for up to 30 days before release. Cohere's framework, written earlier, says nothing about giving government pre-release access.
Security of models, weights and development infrastructure
Cohere runs a defense-in-depth security programme aligned to SOC 2 Type II, strictly restricts access to unreleased weights, secures containers and APIs, and has significant releases independently penetration-tested before deployment: "We cannot secure our models if we do not secure the environment in which we develop them."
Article 55 requires providers of systemic-risk models to ensure an adequate level of cybersecurity protection for the model and its physical infrastructure, which is the ground this part of Cohere's framework covers.
SB 53 requires a large frontier developer's published framework to describe its cybersecurity practices for securing unreleased model weights against unauthorised modification or transfer; Cohere's framework describes exactly such controls.
Harms in context, including discrimination, with duties shared by deployers
Harms are scored for likelihood and severity in the customer's context (its lead example is résumé summarisation for hiring), outputs are tested across languages, guardrails are configurable, and deployers are told they are responsible for "implementing appropriate monitoring controls during deployment".
Annex III lists AI used for recruitment and selection, including analysing and filtering job applications and evaluating candidates, as high-risk, and the Act divides duties between the provider of the high-risk system and its deployer (Article 26), matching Cohere's context-based, shared-responsibility model.
The voluntary NIST AI Risk Management Framework is built on mapping, measuring and managing risk in the context of use, which is Cohere's approach. There is no binding federal counterpart, and America's AI Action Plan directs NIST to revise the framework to remove references to misinformation, diversity, equity and inclusion, and climate change.
Transparency: documentation for customers, government and the public
"Documentation is a key aspect of our accountability to our customers, partners, relevant government agencies, and the wider public"; Cohere publishes model cards, a usage policy, technical guides and a Trust Center, and published this framework.
Article 53 requires every general-purpose model provider to keep technical documentation for the AI Office, supply information to downstream providers and publish a summary of training content. Cohere's model cards and guides serve the downstream-information purpose, but the framework does not mention documentation for regulators or a training-content summary.
SB 53 requires frontier developers to publish a transparency report when they deploy a new or substantially modified frontier model, and large frontier developers to include summaries of their catastrophic-risk assessments, which is more than the model cards the framework commits to.
Reporting serious incidents to authorities
The framework commits to monitoring deployed use where Cohere has visibility and "revoking access from accounts that abuse our systems", and to a responsible disclosure policy for outside researchers, but contains no commitment to report incidents to any government authority.
Article 55 obliges providers of systemic-risk models to track, document and report serious incidents and possible corrective measures to the AI Office and, as appropriate, national authorities without undue delay.
SB 53 requires frontier developers to report critical safety incidents to California's Office of Emergency Services within 15 days, or within 24 hours where there is an imminent risk of death or serious physical injury.
External red teaming, research support and standards bodies
Red teaming "may include independent external parties, such as NIST and Humane Intelligence"; Cohere cooperates with international AI safety institutes, funds outside research through Cohere For AI and a bug bounty, and contributes to OWASP, CoSAI, the Cloud Security Alliance and ML Commons.
The Act relies on harmonised standards and codes of practice drawn up with providers and other stakeholders to make its general-purpose model duties concrete, and requires adversarial testing of systemic-risk models; Cohere's participation in standards work and external red teaming fits that model.
The Action Plan has a section on building an AI evaluations ecosystem and gives NIST's CAISI a lead role in evaluating frontier models, which matches Cohere's practice of red teaming with NIST and submitting models to public benchmarks.
Both the EU and California have since made published risk frameworks a legal matter for the largest developers, and both build them around the severe-risk categories that Cohere's framework declines to treat as its basis. The EU AI Act obliges providers of general-purpose models with systemic risk to evaluate and adversarially test their models, assess and mitigate systemic risks, report serious incidents to the AI Office and secure the model, and the General-Purpose AI Code of Practice, which Cohere says it has signed, asks those providers for a written safety and security framework. California's SB 53 requires large frontier developers to publish a frontier AI framework built around catastrophic-risk thresholds and to report critical safety incidents. Cohere's framework shares the security core of both regimes: defense-in-depth, restricted access to unreleased weights, independent penetration testing, and adversarial testing with outside parties including NIST. It diverges on scope and on reporting. Its release rule is "no significant regressions" against Cohere's own previous model rather than a capability threshold, it limits assurance to risks "known, measurable, or observable today", and it contains no commitment to notify any authority of an incident. Its strongest overlap with the law is in the EU's high-risk regime rather than the frontier rules: its lead example, résumé summarisation for hiring, is a use the AI Act lists as high-risk, and its insistence that deployers carry responsibility for monitoring in their own context matches the Act's split of duties between providers and deployers.
Source
https://cohere.com/security/the-cohere-secure-ai-frontier-model-framework-february-2025.pdf- Date on the page:
- Feb 11, 2025 (announcement page cohere.com/blog/secure-model-framework); no date printed in the PDF
- Source checked:
- opened and confirmed on 2026-09-30