World Economic Forum - AI Governance Framework

AI Agents in Action: Foundations for Evaluation and Governance

World Economic Forum

RAI-XW-GO-EVALUAT-2025
Effective: November 27, 2025
In Force(In Force)
PolicyGovernance and OversightRisk ManagementSafety, Testing, and Evaluation
Export PDF

This WEF framework provides a structured foundation for organizations to evaluate, manage, and govern AI agents responsibly, bridging the gap between adoption and oversight.

Overview

The World Economic Forum, in collaboration with Capgemini, published the white paper "AI Agents in Action: Foundations for Evaluation and Governance" on November 27, 2025. This seminal document addresses the burgeoning challenge of responsibly deploying Artificial Intelligence (AI) agents, which are rapidly transitioning from experimental prototypes to integral collaborators across diverse sectors such as business, public services, and daily life. The report acknowledges that while AI agents offer significant gains in efficiency and new forms of human-machine interaction, their autonomous nature introduces complex risks related to safety, system integration, and trust. It highlights a widening gap between the accelerating pace of AI agent experimentation and the maturity of oversight mechanisms within most organizations. The framework aims to bridge this gap by providing a structured foundation for organizations to evaluate, manage, and govern AI agents effectively, ensuring that their adoption aligns with proportionate safeguards and prepares for increasingly complex multi-agent ecosystems.

The document is primarily tailored for adopters of AI agents, including decision-makers, technical leaders, and practitioners who are looking to integrate these advanced systems into their organizational workflows and services. It emphasizes that as AI agents gain autonomy and decision-making capabilities, organizations must approach their integration with the same rigor as onboarding a new human employee, complete with well-defined roles, safeguards, and structured oversight practices. The white paper is a crucial step in guiding early adopters through the often complex and uneven path of AI agent adoption, advocating for cross-functional efforts and collaborative governance to amplify human ingenuity, promote innovation, and improve overall quality of life. It underscores the necessity of clear standards, continuous monitoring, and scalable governance to foster trustworthy human-AI collaboration and prepare for future agentic systems.

Definitions

The "AI Agents in Action" framework introduces and elaborates on several key terms critical for understanding and governing autonomous AI systems. Central to the document is the definition of "AI Agents" themselves: these are characterized as AI-powered systems capable of independently interpreting information, making decisions, and carrying out actions to achieve specific goals, often with minimal or no continuous human intervention. This distinguishes them from traditional AI models that primarily process inputs and provide outputs, with humans typically making the final decisions. The framework further delves into a "Classification" system for AI agents, which is a structured methodology to articulate an agent's function, role (specialized or generalist), predictability (deterministic or non-deterministic), autonomy (degree of independent planning, decision-making, and action), authority (permissions and system access), use case, and operational environment complexity. This classification is vital for tailoring governance approaches.

Another fundamental concept is "Evaluation" of AI agents, which goes beyond traditional model assessment. It involves examining dimensions such as task success rates, reliability of tool use, performance, behavioral patterns over time, user trust, interaction patterns, and operational robustness within real-world workflows. The framework stresses that real assurance comes from contextual evaluation, meaning testing agents in environments that accurately mirror their actual deployment conditions. Complementing evaluation is "Risk Assessment," which determines the safety and suitability of an AI agent based on the evidence gathered during evaluation. This process identifies potential harms and informs the development of effective mitigation strategies. Finally, the document introduces "Progressive Governance," an adaptive approach where the level of oversight and safeguards applied to an AI agent is directly correlated with its classification and evaluation outcomes. For instance, agents with higher autonomy, greater authority, or those operating in complex, non-deterministic environments require more stringent safeguards and monitoring.

Governance and Institutional Framework

The "AI Agents in Action: Foundations for Evaluation and Governance" white paper is a product of the World Economic Forum's broader AI Governance Alliance, highlighting a collaborative and multi-stakeholder approach to addressing the challenges of AI. The AI Governance Alliance, established by the World Economic Forum, serves as a global multi-stakeholder group that brings together representatives from industry, government, civil society, and academia. Its overarching mission is to promote responsible, inclusive, and accountable AI development and deployment. This collaborative effort is crucial given the cross-cutting nature of AI agent impacts, which necessitate harmonized approaches across various sectors and jurisdictions. The framework itself, developed in collaboration with Capgemini, reflects this multi-stakeholder input, aiming to provide guidance that is both practical for industry adopters and aligned with broader societal values and ethical considerations.

The institutional framework envisioned and supported by this document emphasizes shared responsibility across the AI ecosystem. It outlines the roles of various actors, from AI agent creators and developers to adopters and users, in ensuring the safe and ethical deployment of these systems. The document implicitly positions the AI Governance Alliance as a key facilitator for ongoing dialogue, research, and the co-design of governance protocols. It encourages organizations to establish internal governance structures, such as cross-functional "Agent Councils" (as suggested by related frameworks), to oversee the entire lifecycle of AI agents, from initial design and development through deployment, monitoring, and eventual retirement. This distributed yet coordinated governance model is designed to ensure that accountability, transparency, and risk management are embedded at every stage of an AI agent's operation, rather than being an afterthought.

Key Provisions

The "AI Agents in Action: Foundations for Evaluation and Governance" framework is structured around four foundational pillars designed to guide the responsible adoption and deployment of AI agents. The first key provision is a comprehensive "Classification" system, which helps organizations articulate an agent's characteristics across seven core dimensions: function, role, predictability, autonomy, authority, use case, and environment. This classification is crucial for establishing a shared understanding of an agent's capabilities and operational context, enabling decision-makers to tailor appropriate governance measures. For instance, an "agent card" or "resume" for each AI agent is suggested to provide critical information before its integration into organizational workflows.

The second provision focuses on rigorous "Evaluation," recognizing that AI agents, unlike static models, are dynamic systems. This involves assessing dimensions such as task success rates, tool-use reliability, performance over time, user trust, and operational robustness in real-world settings. The framework advocates for contextual evaluation, mirroring actual deployment conditions, to gain true assurance. The third provision is "Risk Assessment," which translates evaluation evidence into meaningful insights regarding an agent's safety and suitability. This pillar helps identify potential harms and informs the development of robust mitigation strategies. Finally, the document introduces "Progressive Governance," a flexible and adaptive approach that directly links classification and evaluation outcomes to the implementation of proportionate safeguards. This means that agents with higher autonomy, greater authority, or those operating in complex, non-deterministic environments will be subject to more stringent oversight, continuous monitoring, and stricter control mechanisms, ensuring a fit-for-purpose regulatory posture.

Scope and Application

The "AI Agents in Action: Foundations for Evaluation and Governance" white paper is designed to be broadly applicable across various sectors and organizational types that are engaging with or planning to deploy AI agents. Its primary target audience includes adopters of AI agents, encompassing a wide range of stakeholders such as organizational decision-makers, technical leaders, and practitioners. The framework provides a conceptual blueprint for moving AI agents from experimental phases to full-scale deployment within organizational workflows and services. It acknowledges the rapid shift of AI agents from prototypes to integrated collaborators in business, public services, and daily life, recognizing their potential to enhance efficiency and introduce novel human-machine interactions. The document's guidance is particularly relevant for organizations grappling with the unique challenges posed by autonomous systems, including issues of accountability, transparency, and control.

Geographically and jurisdictionally, the framework aims for broad relevance, given the World Economic Forum's international mandate and the global nature of AI development and deployment. While it does not prescribe specific legal mandates, it offers foundational principles and structured approaches that can inform and complement existing or emerging regional and national AI regulations, such as the EU AI Act. The emphasis on a progressive governance approach allows organizations to adapt the framework's recommendations to their specific operational contexts, risk appetites, and the regulatory landscapes in which they operate. By providing a common language and structured methodology for classification, evaluation, risk assessment, and governance, the document facilitates international alignment and best practice sharing, helping to ensure that AI agent adoption is guided by consistent principles of trust, safety, and accountability worldwide.

Implementation Framework

The implementation framework proposed by "AI Agents in Action: Foundations for Evaluation and Governance" is characterized by a progressive and iterative approach, emphasizing the integration of governance practices throughout the entire AI agent lifecycle. Organizations are encouraged to begin with lightweight pilots for low-risk, highly visible workflows to prove value and expose risks in controlled settings, rather than getting bogged down in extensive policy drafting upfront. This practical, learn-by-doing methodology allows for the gradual layering of measurable guardrails and the automation of oversight as the scale and complexity of agent deployment increase. A critical element of implementation involves establishing clear ownership and accountability for every deployed agent, ensuring that a business owner and sponsor are responsible for its objectives, risk profile, and compliance footprint.

Key steps for implementation include mapping the agent lifecycle and its associated risk surface, drafting baseline policies and guardrail requirements, and instrumenting agents for deep observability. This means capturing agent behaviors, decisions, and data interactions through robust logging and monitoring to ensure traceability and meet compliance requirements. Continuous evaluation metrics must be established, and runtime protection and intervention paths need to be implemented to manage unexpected behaviors or failures. The framework also stresses the importance of building an audit-ready documentation and version control trail, as well as operationalizing review and escalation loops to ensure continuous attention to governance. Ultimately, the goal is to scale governance through automation and continuous learning, building a self-reinforcing system that adapts as the technology evolves and as agents operate across organizational boundaries and interact with other systems. This structured approach aims to enable risk-aware scaling, enhance stakeholder trust, and prevent costly setbacks from unmanaged deployments.

Monitoring and Evaluation

The "AI Agents in Action: Foundations for Evaluation and Governance" framework places significant emphasis on robust monitoring and evaluation mechanisms as indispensable components of responsible AI agent deployment. It recognizes that AI agents are dynamic systems that learn and adapt, necessitating continuous oversight rather than one-off assessments. The document advocates for instrumenting agents for deep observability, meaning that their behaviors, decisions, and data interactions must be systematically captured through comprehensive logging and monitoring. This continuous data collection is fundamental for establishing traceability, understanding agent performance in real-world environments, and ensuring compliance with defined policies and ethical standards. Alerts should be triggered for policy breaches and integrated into existing security monitoring dashboards to facilitate timely responses and minimize handoffs between teams.

Evaluation, as a distinct pillar of the framework, involves assessing various dimensions such as task success rates, reliability of tool use, performance over time, user trust, and operational robustness. This goes beyond simple output checks to encompass the entire decision-to-action cycle of an agent. The framework suggests establishing continuous evaluation metrics and conducting contextual evaluations that mirror actual deployment conditions to gain genuine assurance. Furthermore, it highlights the need for systematic processes that make governance feel automatic rather than burdensome, advocating for regular review and escalation loops. For instance, convening an "Agent Council" for focused triage sessions to prune false positives and assign owners for genuine issues is a recommended practice. This iterative monitoring and evaluation approach ensures that organizations can adapt to the evolving capabilities of AI agents, detect misalignments early, and maintain consistency across increasingly complex, multi-agent ecosystems.

Relationship to Other Instruments

The "AI Agents in Action: Foundations for Evaluation and Governance" framework is positioned within a broader landscape of international AI governance instruments and initiatives. While it provides a specific, structured approach for evaluating and governing AI agents, it implicitly or explicitly relates to other global efforts to shape responsible AI development. For instance, the document's emphasis on risk assessment, transparency, accountability, and ethical considerations resonates strongly with principles outlined in documents like the OECD Principles on AI, which promote trustworthy AI, and the European Union's AI Act, which mandates risk-based regulatory requirements for AI systems. The framework's focus on a "shift-left methodology"—integrating safety and ethical considerations early in the development lifecycle—aligns with a growing consensus across many international standards and guidelines that proactive risk management is superior to reactive mitigation.

Furthermore, the World Economic Forum's AI Governance Alliance, which produced this white paper, actively engages with a diverse set of stakeholders including governments, industry, academia, and civil society. This collaborative approach inherently seeks to build bridges and foster interoperability between different regulatory frameworks and best practices globally. The document contributes to the ongoing dialogue within the Centre for the Fourth Industrial Revolution Network, which aims to co-design, test, and refine governance protocols for emerging technologies. By offering a functional classification and a progressive governance approach for AI agents, the framework provides practical guidance that can help organizations navigate the complexities of complying with various national and international regulations, ensuring that their AI agent deployments are not only efficient but also legally compliant, ethically responsible, and operationally safe. It acts as a specialized guide that complements broader AI governance principles by focusing on the unique challenges and opportunities presented by autonomous agentic systems.

International Alignment

The World Economic Forum, as an international NGO, inherently designs its initiatives and publications, including "AI Agents in Action: Foundations for Evaluation and Governance," with a strong emphasis on international alignment and collaboration. The AI Governance Alliance, under whose auspices this white paper was developed, is a global multi-stakeholder group comprising representatives from numerous countries and sectors. This diverse composition ensures that the framework's foundational pillars—classification, evaluation, risk assessment, and progressive governance—are informed by a wide array of perspectives and are designed to be relevant across different legal, cultural, and economic contexts. The document's goal to provide a "structured foundation to close the gap between AI Agents adoption and real-world deployment" is a global imperative, as AI agents are being developed and deployed across borders, necessitating harmonized approaches to ensure safety and trust.

The framework's principles, such as promoting transparency, accountability, and the integration of safeguards, are consistent with broader international efforts to establish trustworthy AI, including those advanced by organizations like the OECD, G7, and the Council of Europe. By offering a common language and methodology for understanding and managing AI agents, the document facilitates cross-border cooperation and the mutual recognition of best practices. It helps to prevent regulatory fragmentation by providing a flexible structure that can be adapted to various national regulatory regimes while promoting a consistent baseline for responsible AI agent development and deployment. The World Economic Forum's role in convening global leaders and experts through platforms like the Centre for the Fourth Industrial Revolution further underscores its commitment to fostering international dialogue and consensus-building on critical technological governance issues, ensuring that frameworks like "AI Agents in Action" contribute to a globally coherent and responsible AI future.

Implementation Timeline

MilestoneDateStatus
Publication of White Paper2025-11-27In Force
Integration of AI Agents (82% of executives planning)Within 1-3 years of publication (by 2028)Ongoing / Planned
Development of multi-agent ecosystemsOngoing / FutureEmerging

Adoption and Endorsement

EntityDateStatus
World Economic Forum2025-11-27Adopted
Capgemini (in collaboration)2025-11-27Endorsed
AI Governance Alliance (WEF)2025-11-27Adopted
Early Adopters (decision-makers, technical leaders, practitioners)OngoingImplementing

Sources and References

SourceType
AI Agents in Action: Foundations for Evaluation and Governanceofficial
Plain English

The World Economic Forum's "AI Agents in Action" framework, published November 27, 2025, offers a structured guide for organizations to responsibly evaluate, manage, and govern their AI agents. This framework applies to any organization, from decision-makers to technical teams, looking to integrate autonomous AI agents into their business, public services, or daily operations. It's designed for early adopters navigating the rapid shift of AI agents from experimental prototypes to integral collaborators.

The framework outlines four key pillars for responsible deployment. Organizations must: - **Classify** each AI agent by its function, autonomy, authority, and operational environment, creating a clear "agent card" or "resume." - **Evaluate** agents continuously, assessing their task success, reliability, performance, and user trust in real-world conditions. - Conduct thorough **risk assessments** to identify potential harms and develop mitigation strategies. - Implement **progressive governance**, meaning the level of oversight and safeguards should directly match the agent's complexity and autonomy – more powerful agents require stricter controls and continuous monitoring.

This framework became effective upon its publication on November 27, 2025. As a World Economic Forum policy, it doesn't carry direct legal penalties. Instead, its "teeth" come from guiding best practices to prevent costly operational setbacks, reputational damage, and potential non-compliance with future national or regional AI laws that may draw from these principles. A key practical pitfall to avoid is getting stuck in extensive policy drafting upfront. The framework encourages starting with lightweight pilots for low-risk workflows, gradually adding measurable safeguards and automating oversight as deployments scale. This "learn-by-doing" approach is crucial for adapting to evolving AI agent capabilities.

Plain-English rewrite by Regulations.ai — not legal advice. Verify against the official text.

What you must do — compliance checklist

0 / 15 marked complete

Plain-English obligations under World Economic Forum - AI Governance Framework. Not legal advice — verify against the official text before relying on it.

  1. #1ImportantOverview

    Applies to: Adopters of AI agents

    organizations must approach their integration with the same rigor as onboarding a new human employee, complete with well-defined roles, safeguards, and structured oversight practices.
  2. #2ImportantImplementation Framework

    Applies to: Organizations deploying AI agents

    establishing clear ownership and accountability for every deployed agent, ensuring a business owner and sponsor are responsible for its objectives, risk profile, and compliance.
  3. #3ImportantImplementation Framework, Monitoring and Evaluation

    Applies to: Organizations deploying AI agents

    instrumenting agents for deep observability. This means capturing agent behaviors, decisions, and data interactions through robust logging and monitoring to ensure traceability.
  4. #4ImportantImplementation Framework, Monitoring and Evaluation

    Applies to: Organizations deploying AI agents

    Continuous evaluation metrics must be established, and runtime protection and intervention paths need to be implemented to manage unexpected behaviors or failures.
  5. #5ImportantImplementation Framework

    Applies to: Organizations deploying AI agents

    runtime protection and intervention paths need to be implemented to manage unexpected behaviors or failures.
  6. #6ImportantImplementation Framework

    Applies to: Organizations deploying AI agents

    building an audit-ready documentation and version control trail, as well as operationalizing review and escalation loops to ensure continuous attention to governance.
  7. #7ImportantImplementation Framework

    Applies to: Organizations deploying AI agents

    operationalizing review and escalation loops to ensure continuous attention to governance.
  8. #8ImportantKey Provisions, Definitions

    Applies to: Adopters of AI agents

    The first key provision is a comprehensive "Classification" system, which helps organizations articulate an agent's characteristics across seven core dimensions.
  9. #9ImportantKey Provisions, Definitions

    Applies to: Adopters of AI agents

    The second provision focuses on rigorous "Evaluation," recognizing that AI agents, unlike static models, are dynamic systems.
  10. #10ImportantKey Provisions, Definitions

    Applies to: Adopters of AI agents

    The third provision is "Risk Assessment," which translates evaluation evidence into meaningful insights regarding an agent's safety and suitability.
  11. #11ImportantKey Provisions, Definitions

    Applies to: Adopters of AI agents

    Finally, the document introduces "Progressive Governance," a flexible and adaptive approach that directly links classification and evaluation outcomes to the implementation of proportionate safeguards.
  12. #12ImportantImplementation Framework

    Applies to: Organizations deploying AI agents

    Key steps for implementation include mapping the agent lifecycle and its associated risk surface, drafting baseline policies and guardrail requirements.
  13. #13RecommendedMonitoring and Evaluation

    Applies to: Organizations deploying AI agents

    Alerts should be triggered for policy breaches and integrated into existing security monitoring dashboards to facilitate timely responses.
  14. #14RecommendedGovernance and Institutional Framework, Monitoring and Evaluation

    Applies to: Organizations deploying AI agents

    It encourages organizations to establish internal governance structures, such as cross-functional "Agent Councils" (as suggested by related frameworks), to oversee the entire lifecycle of AI agents.
  15. #15RecommendedImplementation Framework

    Applies to: Organizations deploying AI agents

    Organizations are encouraged to begin with lightweight pilots for low-risk, highly visible workflows to prove value and expose risks in controlled settings.

© Regulations.AI — created on 09-Jan-2026 using Gemini 2.5 Flash