Evaluation
Systematic assessment of AI products against set criteria.
Definitions (2)
The systematic assessment process by which AI products and systems are tested and measured against the AIEC's established criteria and guidelines across multiple dimensions (e.g., safety, fairness, privacy, robustness) to determine compliance and risk classification.
The process of examining an AI agent's performance, behavior over time, reliability, user trust, and operational robustness in dynamic, real-world environments.
Related Terms
AI Evaluation
Systematic assessment of AI products for safety, fairness, and compliance....
Risk assessment
Procedure to identify and mitigate social, ethical and legal risks....
AI Impact Assessment
Documented risk and harms analysis before AI deployment....
Evaluation of Impact
Human-rights-focused impact assessment (DPIA-like) for AI projects....
evaluations
Empirical tests and analyses measuring capabilities and potential harms....