Large Language Model
A machine-learning model trained on very large text datasets to understand and generate human language across many tasks.
Definition
Official / legal framing: In technical standards a large language model (LLM) is defined as a machine‑learning model that encodes the functioning of natural language with a large number of parameters and that facilitates a variety of natural language processing (NLP) tasks (e.g., generation, summarization, translation, classification). This formulation is set out in ISO/IEC 22989 (AI concepts and terminology) which describes an LLM as a machine learning model that “encodes the functioning of natural language… with a large number of parameters” and notes their use in many NLP tasks. ([standards.iteh.ai](https://standards.iteh.ai/catalog/standards/cen/27af3d68-cb4d-4873-a622-bdf3b5ca8602/en-iso-iec-22989-2023-pra1-2025?utm_source=openai))
Jurisdictional variations:
- European Union: The EU AI Act does not use the exact term “large language model” as a standalone defined item but regulates the class of general‑purpose AI models (GPAI models / foundation models) that explicitly encompasses LLMs: "an AI model, including where such an AI model is trained with a large amount of data using self‑supervision at scale, that displays significant generality…" (AI Act, Article 3(63)). Providers of such models are subject to bespoke obligations under the Act. ([ai-act-service-desk.ec.europa.eu](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-3?utm_source=openai))
- United States (federal): Executive Order No. 14110 (and implementing NIST guidance) treats LLMs as a subset of "foundation models" or "generative AI" and, for policy purposes, distinguishes "dual‑use foundation models" by capability (e.g., "tens of billions of parameters") and by potential for serious national‑security or public‑safety risk; NIST’s Generative AI Profile (AI RMF companion) further operationalizes risk and mitigation guidance for large models. ([presidency.ucsb.edu](https://www.presidency.ucsb.edu/documents/executive-order-14110-safe-secure-and-trustworthy-development-and-use-artificial?utm_source=openai))
- United States (state level): Some state laws regulate uses of AI systems generally (e.g., Colorado SB24‑205 regulates "high‑risk AI systems" and places obligations on developers and deployers) but do not generally supply a standalone canonical statutory definition of "LLM"; instead states typically regulate by function, risk, or use case. ([s3.us-west-2.amazonaws.com](https://s3.us-west-2.amazonaws.com/beta.leg.colorado.gov/8ae60739b2b5dac9235add08baebc925))
- International / standards bodies: ISO/IEC 22989 provides an explicit technical definition (quoted above) and the OECD, UNESCO and other international instruments treat LLMs as examples of foundation or generative models and address governance principles (transparency, data governance, human oversight). ([standards.iteh.ai](https://standards.iteh.ai/catalog/standards/cen/27af3d68-cb4d-4873-a622-bdf3b5ca8602/en-iso-iec-22989-2023-pra1-2025?utm_source=openai))
Context and scope: LLMs are a subset of foundation models focused on text. Typical technical features include very large parameter counts (often billions to trillions), pretraining on massive text corpora (frequently via self‑supervised objectives), transformer or related architectures, and the ability to be fine‑tuned or prompted for many downstream NLP tasks. They are commonly used as components inside broader AI systems (chatbots, search, summarizers, code assistants) and therefore are regulated either as models (when obligations target model providers) or as systems/applications (when obligations target deployers). See ISO/IEC 22989 for the technical definition and the EU AI Act for the regulatory categorisation of general‑purpose/foundation models. ([standards.iteh.ai](https://standards.iteh.ai/catalog/standards/cen/27af3d68-cb4d-4873-a622-bdf3b5ca8602/en-iso-iec-22989-2023-pra1-2025?utm_source=openai))
Practical implications for businesses:
- Classification & compliance: LLMs will often fall within the EU AI Act’s "general‑purpose AI model" regime (bringing provider transparency, risk‑management and information obligations) and under US/NIST guidance for GAI safety and documentation; businesses must assess whether they are a developer/provider, deployer, importer or downstream integrator and comply accordingly. ([ai-act-service-desk.ec.europa.eu](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-3?utm_source=openai))
- Data and IP risk: training on large scraped or licensed corpora raises privacy, copyright and data‑use risks—FTC guidance emphasises that using data contrary to privacy promises can trigger enforcement. Documentation of training data provenance and licences is therefore a practical necessity. ([ftc.gov](https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments?utm_source=openai))
- Operational controls: regulators and standards expect measures such as red‑teaming, pre‑deployment testing, incident reporting, and post‑market monitoring for large models or systems built from them. The EU AI Act and NIST GAI profile both highlight testing, transparency, and incident management as core expectations. ([ai-act-service-desk.ec.europa.eu](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-3?utm_source=openai))
Key criteria & requirements (typical):
Authorities and standards use different signals to identify "large" or "foundation" models; these signals are not uniform but commonly include:
- Generality of capability: ability to perform a wide range of tasks (EU AI Act: general‑purpose AI model). ([ai-act-service-desk.ec.europa.eu](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-3?utm_source=openai))
- Training scale & method: trained on broad datasets, often via self‑supervision. ([standards.iteh.ai](https://standards.iteh.ai/catalog/standards/cen/27af3d68-cb4d-4873-a622-bdf3b5ca8602/en-iso-iec-22989-2023-pra1-2025?utm_source=openai))
- Model size / compute: number of parameters and/or amount of training compute (EO 14110 and some policy texts use numeric thresholds such as "tens of billions of parameters" or compute‑based tests for reporting purposes, though such thresholds are policy choices and may be updated). ([presidency.ucsb.edu](https://www.presidency.ucsb.edu/documents/executive-order-14110-safe-secure-and-trustworthy-development-and-use-artificial?utm_source=openai))
- Risk profile and intended use: if the model or system can cause serious harms (safety, public security, large‑scale misinformation, etc.), additional obligations (reporting, red teaming, incident disclosure) are likely. ([ai-act-service-desk.ec.europa.eu](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-3?utm_source=openai))
Examples & cross‑references: Representative LLM‑style models in public discourse include families such as OpenAI’s GPT models, Google/DeepMind’s PaLM/Gemini family, Meta’s LLaMA derivatives and other large transformer‑based language models; these examples illustrate how LLMs operate as foundation components that can be integrated into conversational agents or domain applications. Related regulatory and technical concepts to review include foundation model / general‑purpose AI model (EU AI Act), generative AI, model card / system card, training data provenance, and red‑teaming / pre‑deployment testing. ([ai-act-service-desk.ec.europa.eu](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-3?utm_source=openai))
Sources
- •NIST AI RMF
- •Technical Literature
Related Terms
Foundation Model
A large AI model trained on broad data that is designed to be adapted for many different tasks (e.g., LLMs, multimodal models)....
General-Purpose AI Model
An AI model trained at scale that displays significant generality and can competently perform many distinct tasks and be integrated into varied downstream systems....
Generative AI
A class of AI systems that produce new digital content (text, image, audio, video, code) by modelling and emulating patterns in training data....
Training Data
Data used to develop and train AI models to perform their intended functions....
Model Card
Standardized documentation providing key information about an AI model....