A Proposal to Identify High-Impact Capabilities of General-Purpose AI Models
Hobbhahn, M., Hovy, D., Vanschoren, J., Fernandez Llorca, D., Eriksson, M., Gomez, E.
Hobbhahn, M. et al. - European Commission, Joint Research Centre (Publications Office of the European Union), 2025-10-10
Abstract
This report proposes a scientific methodology to identify high-impact capabilities in General-Purpose AI (GPAI) models, defined in the EU AI Act as capabilities of the most advanced GPAI models. High-impact capabilities play an important role in the EU AI Act since GPAI models with high-impact capabilities are classified as GPAI models with systemic risks. The approach is based on observational scaling laws using Principal Components Analysis (PCA) from a set of existing benchmarks, allowing for the extraction of a low-dimensional capability measure that can be used to identify models with high-impact capabilities. The proposed method involves selecting a diverse set of benchmarks that measure general capabilities, such as MMLU-Pro, GPQA-diamond, MATH-level-5, and HumanEval, and aggregating their scores using a weighted threshold-based metric. The weights are determined by the PCA approach, and the threshold is based on a reference model, to be set by the enforcement authority based on legal, policy, and risks considerations.