Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control
Independent International Scientific Panel on AI
Independent International Scientific Panel on AI - United Nations, 2026-09-21
Abstract
The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions. Between May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.