To Steer the AI Frontier, Washington Must Build Its Testing Power
Published September 25, 2026 · Printed in the byline under the title ('By Francis deSouza · September 25, 2026 · 4 min read') and matched by the page's datePublished metadata (2026-09-25). The post refers to 'this week's UN General Assembly', which fits that date.
Not law. This is a company's own public position on AI regulation. It is not law, and it carries no legal force.
What it argues for
This is a short signed essay on the scale.com blog by Francis deSouza, published the week the UN General Assembly discussed AI, and it is now Scale AI's newest statement of what it wants from government. Its argument is that AI policy should be driven by measured evidence rather than by warnings: "Much of today's debate about AI risk rests on speculation." It says the costs of error run both ways, since "Moving recklessly would be a mistake, but slowing down out of fear would be just as wrong." Its central claim is that "The federal government's ability to evaluate frontier models has not kept pace with the companies building them", and that on national security and public safety "Washington cannot let the labs grade their own homework." It asks for three things: an accountable official to coordinate testing across federal agencies, a stronger government ability to evaluate frontier models before deployment, secured by working with developers, and sustained public funding for independent public-private testing. It calls the proposed budgets for CAISI, the government's frontier AI testing center, inadequate. It then offers Scale's own evaluation work and research as evidence of what testing finds: agents that can be hijacked through their clarifying questions, and models that refuse to help cyber defenders. This marks a clear move from Scale's March 2025 submission on the AI Action Plan. That document asked for a regulatory framework that governs uses of AI rather than the technology, and treated measurement mainly as a field for US standards leadership. This essay instead asks the federal government itself to examine frontier models before they are released. It does not say whether that access should be compulsory, and it leaves out the export, defence-adoption and workforce agenda of the earlier document. It closes: "The nation that can measure the frontier will be the one that steers it."
Stated positions (15)
- Policy should rest on evidence, not warnings: "Some of the loudest warnings have little grounding in science, yet they drive policy as if they were proven." It says "The way forward is evidence."
- Errors cost in both directions: "Moving recklessly would be a mistake, but slowing down out of fear would be just as wrong." It counts delay as a loss, since "Every week lost is a week of slower progress toward new cancer treatments, stronger American national security, and scientific breakthroughs."
- International agreement needs a testing base: of the UN General Assembly talks it says "Any agreement they reach will only be as strong as the evidence beneath it", and that governments must settle "what are the benchmarks, who verifies the results, and what happens when serious risks are found."
- Government testing capacity is behind: "The federal government's ability to evaluate frontier models has not kept pace with the companies building them. That gap needs to close." In short, "You cannot regulate what you cannot measure."
- Testing must be independent of the labs: "Policymakers need independent testing to reliably separate demonstrated risks from imagined ones." It grants that labs have "a moral responsibility and a strong commercial incentive to test their systems", but says "Washington cannot let the labs grade their own homework."
- Government needs its own evidence on named risk areas: "The government needs its own evidence on cyber, biological, autonomy-related, and other emerging risks."
- A clear mandate: "Designate an accountable official to coordinate testing across federal agencies, set government-wide evaluation priorities, mobilize evaluation expertise across the public and private sectors, and ensure serious risks are identified and mitigated."
- Pre-deployment evaluation: "Strengthen the government's ability to evaluate frontier models before deployment by working with model developers to secure access for testing and reassess national security risks as capabilities evolve and new uses emerge." It does not say whether developers should be required to give that access.
- Dedicated capacity: "Commit sustained government funding to independent public-private testing, including the testing tools, resources, and technical expertise needed across national security domains."
- CAISI is underfunded: it notes that "The President requested $27 million for CAISI" for fiscal year 2027 "while House appropriators proposed up to $15 million", and says "these funding levels are like managing JFK air traffic control from a small-town control tower."
- Government should draw on industry rather than build alone, and Scale offers itself: "At Scale, we've spent years developing the expertise and building the infrastructure to test advanced models on the hardest problems." It says it has partnered with government AI evaluation bodies in the United States, the United Kingdom, Korea and Singapore.
- Standard tests miss real attacks. Scale's research found that an attacker can plant instructions in the answer to an agent's clarifying question: "Most models tested were vulnerable to this form of attack." For several models such attacks "succeeded about 10 times more often" than instructions hidden in emails, documents or webpages.
- Over-refusal is also a risk: in a national cyber defence competition, "AI models refused to help the defenders 12 percent of the time, and more than 40 percent of the time when they asked how to lock down their systems."
- Testing should prepare critical sectors: "Without testing built for these scenarios, policymakers cannot know what America's banks, power grids, and hospitals need to be ready for."
- The aim is to act before harm and then clear the way: with the capacity to direct outside experts, Washington "can address real risks before they become crises and clear the way for everything AI promises." It ends: "The nation that can measure the frontier will be the one that steers it."
About this document
A short blog post of about 800 words in the Company section of the scale.com blog, headed 'To Steer the AI Frontier, Washington Must Build Its Testing Power' and signed by Francis deSouza; a site-wide banner on the same page announces his appointment as Scale's new CEO. It has no section headings. It runs as prose: the case for evidence over speculation, a paragraph on the UN General Assembly, the federal testing gap, then three labelled asks ('A clear mandate', 'Pre-deployment evaluation', 'Dedicated capacity'). A paragraph on CAISI funding follows, then a section on Scale's own evaluation partnerships and cyber research (agent hijacking through clarifying questions, and model refusals in a national cyber defence competition), and a short close. It links to the NIST budget submission, a House appropriations report and Scale's research, and names no law or bill.
How this sits against AI law
Each stance compared with what EU and US instruments actually require. Where no instrument addresses a theme, that gap is shown rather than hidden.
One accountable federal official to coordinate AI testing
The government needs a clear mandate: an accountable official who coordinates testing across federal agencies, sets government-wide evaluation priorities, mobilises public and private evaluation expertise, and ensures "serious risks are identified and mitigated."
The AI Act already gives the EU a single accountable body for frontier models: the Commission's AI Office supervises and enforces the general-purpose AI obligations and can itself evaluate general-purpose models with systemic risk (Article 92). That body also has enforcement powers, which Scale does not ask for.
Executive Order 14409 spreads frontier-model assessment across several agencies. Section 3 tasks Treasury, the NSA and CISA, in consultation with the National Cyber Director, the President's science adviser and NIST, and gives the 'covered frontier model' determination to the NSA Director. It names no single official to set government-wide evaluation priorities, so Scale asks for more.
Government evaluation of frontier models before deployment
The government should strengthen its ability to evaluate frontier models before deployment, "by working with model developers to secure access for testing" and reassessing national security risks as capabilities and uses change. Scale does not say whether that access should be compulsory.
The Act makes evaluation binding rather than cooperative. Article 55 requires providers of general-purpose models with systemic risk to evaluate them, including adversarial testing, and to report serious incidents, and Article 92 lets the AI Office evaluate such models itself and request access to them. Scale asks only for access negotiated with developers, although the Act sets no dedicated government test before release.
Section 3 of Executive Order 14409 has agencies design a voluntary framework through which developers can give the government access to 'covered frontier models' for up to 30 days before releasing them to other trusted partners, after a classified benchmarking of cyber capabilities. Section 3(c) rules out any mandatory licensing, preclearance or permitting. That is the cooperative pre-deployment access Scale describes.
Government's own evidence on national-security risks
Labs should not "grade their own homework" on national security and public safety; "The government needs its own evidence on cyber, biological, autonomy-related, and other emerging risks."
The EU's General-Purpose AI Code of Practice, whose Safety and Security chapter applies to systemic-risk models, has signatory providers assess and mitigate systemic risks such as CBRN, cyber-offence and loss-of-control risks and report to the AI Office. The assessment is still the provider's own, so Scale's demand for separate government evidence goes further than the Code.
The Action Plan's section 'Ensure that the U.S. Government is at the Forefront of Evaluating National Security Risks in Frontier Models' has CAISI evaluate frontier systems for national security risks with agencies expert in CBRNE and cyber risks. It also has CAISI build and maintain national-security evaluations with national security agencies and research institutions.
Sustained public funding for independent testing capacity
Government should "Commit sustained government funding to independent public-private testing", and current CAISI budgets ($27 million requested for fiscal year 2027, up to $15 million proposed by House appropriators) are far too small.
The Act creates the AI Office, a scientific panel of independent experts (Article 68) and Union AI testing support structures (Article 84), but as a regulation it sets no budget for testing capacity. Funding is decided in the EU budget, not in the Act.
The March 2026 legislative recommendations ask Congress to "ensure that the appropriate agencies within the national security enterprise possess sufficient technical capacity to understand frontier AI model capabilities", in consultation with frontier developers. They name no funding level and no testing body, and Scale calls the budgets now proposed for CAISI inadequate.
Testing for real-world cyber risks, including agent hijacking and refusals to help defenders
Standard tests miss how attacks work in practice: agents can be hijacked through the answers to their clarifying questions, and models refuse to help defenders secure their systems. Without testing built for such scenarios, policymakers cannot know what banks, power grids and hospitals must prepare for.
Article 15 requires high-risk AI systems to be resilient against attempts by unauthorised third parties to alter their use or outputs by exploiting vulnerabilities, and Article 55 requires adversarial testing of systemic-risk models. That covers the attack side. Nothing in the Act addresses models refusing legitimate defensive work.
Executive Order 14409 moves on the defenders' side. Section 2 orders agencies to widen access to AI-enabled cyber-defence tools, including covered frontier models for critical-infrastructure operators such as rural hospitals, community banks and local utilities, and to form an AI cybersecurity clearinghouse. It does not call for testing of agent hijacking or of model refusals.
Against the United States this essay mostly asks for more of what federal policy already sets out. The July 2025 Action Plan already has CAISI evaluate frontier systems for national security risks in partnership with developers. Executive Order 14409 of June 2026 already sets up a voluntary route for developers to give the government access to the most cyber-capable models up to 30 days before wider release, and rules out mandatory licensing or preclearance. Scale's ask for pre-deployment access "by working with model developers" fits that voluntary model. Where Scale goes further is on structure and money: one accountable official rather than duties split across Treasury, the NSA, CISA and NIST, and more funding than the budgets now proposed. Against the EU the picture reverses. The AI Act already binds providers of systemic-risk general-purpose models to evaluate and adversarially test them and report serious incidents, and gives the AI Office its own powers to evaluate such models. The EU therefore compels more than Scale asks for, though it has no dedicated pre-deployment government test. This is also a shift for Scale itself. Its March 2025 submission argued that regulation should govern uses of AI rather than the technology, and this essay asks the government to examine the models themselves before release.
Source
https://scale.com/blog/steer-the-ai-frontier-washington-must-build-its-testing-power- Date on the page:
- September 25, 2026
- Source checked:
- opened and confirmed on 2026-09-30