Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control

Independent International Scientific Panel on AI

Independent International Scientific Panel on AI - United Nations, 2026-09-21

0 citations2026

Abstract

The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions. Between May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.

Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control - Research - Regulations.AI | Regulations.ai