Alignment
Ensuring an AI system’s objectives, outputs and behaviour are consistent with specified human goals, values, laws and expectations.
Definition
Techniques and interventions applied to foundation models to ensure outputs are safe, reliable, and conform to stakeholder values or policy objectives; this includes training, reinforcement, filtering, or other safety-oriented processes to mitigate harms such as hallucinations or misleading outputs.
Related Terms
Human-in-the-Loop
An oversight model where humans actively participate in every AI decision cycle, approving actions before execution....
Risk Management
Systematic process of managing AI-related risks....
Explainability
The ability to understand and articulate how an AI system reaches its decisions....
Impact Assessment
A systematic, documented review (ex‑ante or ongoing) of how an AI system may affect individuals, groups, organisations or society and measures to mitigate harms....