Technical

Alignment

Ensuring an AI system’s objectives, outputs and behaviour are consistent with specified human goals, values, laws and expectations.

Definition

Techniques and interventions applied to foundation models to ensure outputs are safe, reliable, and conform to stakeholder values or policy objectives; this includes training, reinforcement, filtering, or other safety-oriented processes to mitigate harms such as hallucinations or misleading outputs.