Prompt Injection
Attack technique where malicious instructions are inserted into inputs to manipulate AI system behavior.
Definition
An adversarial technique in which crafted inputs (prompts) are used to cause an AI system—particularly generative models—to reveal sensitive data, bypass safeguards, or execute unintended actions; the Playbook treats prompt injection as a security threat requiring testing, input sanitisation and contract/supplier controls.
Related Terms
Cybersecurity Requirements
Technical and organisational measures that ensure AI systems resist, detect, respond to, and recover from cyber threats across their lifecycle....
Guardrails
Technical mechanisms and constraints built into AI systems to prevent harmful outputs or behaviors....
Red Teaming
A structured, adversarial testing process that simulates intentional misuse to find vulnerabilities, failure modes, and harms in AI systems before deployment....