Frontier Models

RawGraph

Frontier models are artificial intelligence models at or near a selected boundary of capability, scale, or risk. The label is relational rather than permanent: it compares a model with a reference set, an evaluation suite, or a policy trigger at a particular time. A model described as frontier in one domain may not lead in another, and a model at the research frontier can later become ordinary. There is no universal scientific test, fixed compute threshold, or authoritative current roster that decides frontier status for every purpose.

The term has at least three uses. In descriptive technical writing, it is shorthand for models near the current performance boundary. In safety and governance work, it narrows attention to unusually capable general-purpose models or systems whose capabilities may create severe risks. In law, it means only what a particular instrument defines. These uses overlap, but they are not interchangeable. The European Union, for example, regulates "general-purpose AI models with systemic risk" rather than adopting frontier model as its legal category, while California and New York statutes use their own compute-based definitions.[1][2][8][13][14]

Most systems called frontier models are foundation models, and many are large language models, but neither relationship is definitional. Foundation model describes a broad model trained on broad data and adaptable to many downstream tasks. Frontier describes a changing position or a policy-selected risk class. A model is also not the same thing as a deployed system: prompts, retrieval, tools, memory, access controls, inference resources, and user context can change the capabilities and risks observed in use.[4][5][6]

Vagueness in the word does not mean there is nothing concrete to state. As of August 2026 at least three binding instruments define a covered class by a numeric training-compute threshold, four large developers publish safety frameworks that are triggered mainly by capability that decide what they will and will not release, and several independent groups publish dated capability and compute measurements. Because the underlying population and evaluation methods change quickly, durable reporting identifies the definition, date, comparison set, model or system configuration, and evidence behind a frontier claim. This article therefore explains how the term is used and assessed, and states the specific instruments and measurements that give it content, rather than maintaining a leaderboard of current products.

Meaning and scope

The frontier metaphor refers to an outer boundary. For AI, the boundary can be drawn over several dimensions, such as software engineering, scientific reasoning, multimodal perception, autonomy, cybersecurity, or biological knowledge. Those dimensions need not move together. The 2026 International AI Safety Report describes leading general-purpose AI capabilities as "jagged": a system can perform strongly on some difficult tasks while failing on apparently simpler ones.[6] A single rank or average score therefore hides which boundary is being measured.

Four recurring meanings should be kept separate:

UseWhat the label identifiesWhat must be stated
Descriptive frontierModels near the best demonstrated capability under a chosen comparisonDomain, benchmark or task, access conditions, comparison set, and date
Technical frontierA moving empirical boundary studied through compute, data, algorithms, or evaluationsMeasurement method, model or system configuration, and uncertainty
Policy or risk frontierModels or systems selected for heightened assessment because of potentially severe capabilities or impactsThreat model, risk threshold, safeguards, and responsible actor
Legal frontierA class created by a statute, regulation, order, or guidance documentJurisdiction, operative text, effective date, threshold, exemptions, and legal status

The UK discussion paper prepared for the 2023 AI Safety Summit used frontier AI for highly capable general-purpose models that matched or exceeded the most advanced models then available. It expressly limited that definition to the summit and stated that the paper was not a policy position of the UK government.[1] The Bletchley Declaration used a related but broader formulation, covering highly capable general-purpose models, foundation models, and relevant narrow AI capable of harm.[2] These are influential policy descriptions, not a scientific standard for labeling products.

An academic policy proposal published in 2023 used a risk-oriented definition: highly capable foundation models that could possess dangerous capabilities sufficient to pose severe public-safety risks. Its authors organized the regulatory problem around unpredictable capability development, difficulty preventing misuse, and proliferation.[3] That proposal helped establish a governance vocabulary, but its definition should not be presented as a settled technical consensus or as binding law.

A foundation model is a model trained on broad data that can be adapted to a wide range of downstream tasks. The term emphasizes the model's role as a reusable base for applications.[4] A foundation model can be far from the current capability frontier, and a frontier system need not be a language model.

A general-purpose AI model is a separate functional or legal concept. Under the EU AI Act, it is a model with significant generality that can competently perform a wide range of distinct tasks and can be integrated into varied downstream systems.[8] The Commission's later guidance offers technical indicators for interpreting that definition, but the guidance is nonbinding and still requires case-by-case analysis.[9]

An AI system combines one or more models with other technical and organizational components. The OECD describes a model as a core component used to make inferences, while a system may also include objectives, interfaces, software, data flows, and deployment conditions.[5] An AI agent, for example, may add planning loops, tools, memory, permissions, and environmental feedback around a model. Claims about an agent's ability cannot automatically be attributed to the underlying weights alone.

"Advanced AI," "highly capable AI," and "frontier AI" are often used as umbrella terms in international discussions. They should be treated as source-specific language unless the source supplies a definition.

A moving technical frontier

Frontier status changes for two reasons. First, new models and systems improve or extend the set of demonstrated capabilities. Second, the measurement process changes: evaluations are revised, prompts and tools improve, older test data become contaminated, and researchers learn how to elicit capabilities that an initial test missed. The reference class can also change. "Best publicly documented model," "best model tested by a regulator," and "best commercially accessible system" may refer to different sets.

The boundary is multi-dimensional. A useful technical claim names a capability vector rather than treating intelligence as one scalar. Relevant dimensions can include:

  • task performance and error severity;
  • robustness to distribution shift or adversarial inputs;
  • reliability over long sequences of actions;
  • ability to use tools or operate with limited supervision;
  • potentially hazardous capability in a stated domain;
  • efficiency, latency, and resource requirements;
  • language, modality, and accessibility coverage.

The benchmark chosen for one dimension is not automatically a measure of another. Accuracy on a science question set does not establish safe laboratory assistance, and success in a constrained coding environment does not establish reliable software work in an unfamiliar production repository. Holistic evaluation research has therefore argued for standardized scenarios and multiple metrics, including calibration, robustness, fairness, toxicity, and efficiency rather than accuracy alone.[17]

The unit of comparison matters as much as the metric. A base model, an instruction-tuned model, an API service, and a tool-using system may share weights but behave differently. System prompts, sampling settings, retrieval sources, inference-time computation, tool permissions, and safety filters can all alter measured performance. A frontier claim should identify the tested version and configuration instead of using a product family name as if it were one stable object.

How fast the boundary has actually moved

Attribution matters for these figures. The most systematic public tracking comes from Epoch AI, which maintains a database of notable models and publishes trend estimates with explicit confidence ratings. Epoch classifies individual compute records as confident (within a factor of 3), likely (within a factor of 10), or speculative (within a factor of 30), so single-model numbers should be read as estimates rather than disclosures.

MeasurementEpoch AI's stated findingAs of
Training compute of frontier language modelsGrowing about 5 times per year since 2020, a doubling roughly every 5.2 months; the top-5 trend has risen by a factor of about 10,000 since 2020Updated 5 February 2026[27]
Models trained above 10^25 FLOP33 identified; only two then estimated above 10^26 FLOP, Grok 3 at about 4.6x10^26 and Claude Opus 4 at about 1.5x10^26Updated 6 June 2025[28]
Capability progress on the Epoch Capabilities IndexAbout 8.3 index points per year before April 2024 and about 15.5 after, an acceleration of roughly 1.85 times across 149 modelsAnalysis to December 2025[31]
Turnover at the top17 models have held the highest index score since March 2023; GPT-4 held it for 352 days, about 3.6 times longer than the next longest run2 July 2026[32]
Compute in a single data centerRecord capacity doubling about every 7 months, roughly 3.3 times per year; the tracked sample covered 5.2 million operational H100-equivalents, about 26 percent of global capacity11 June 2026[33]

Two things follow. The input frontier has moved by roughly four orders of magnitude in six years, which is why a fixed statutory compute number ages quickly. And the tenure of any single leading model has collapsed from about a year to a few months, which is why a current-model roster is a poor definition of a durable class.

Compute as an observable proxy

Training compute is attractive to policymakers because it can be estimated before deployment, is less dependent on a particular benchmark, and historically grew rapidly among notable machine-learning systems. A peer-reviewed study of compute trends found a distinct large-scale era in which selected milestone models used substantially more training computation than the preceding trend.[7] Compute is also relevant to infrastructure, concentration, security, and compute governance.

Compute is not a direct measure of generality, capability, or risk. The same operation count can produce different results because of data quality, architecture, optimization, hardware utilization, and algorithmic efficiency. Pretraining compute also excludes or incompletely represents fine-tuning, reinforcement learning, inference-time search, tools, and system scaffolding. Conversely, a costly run can fail or yield a model that does not lead on the capability of interest. Empirical scaling laws describe observed regularities under stated conditions; they are not guarantees that a given amount of compute produces a given behavior.[6][7]

Measurement practice reinforces the point. Epoch AI reports that benchmark-based compute estimates land within about 2.2 times the true value on average, with a 90 percent interval running from 1.1 to 5.0 times, and describes such figures as fairly speculative.[28] A threshold expressed to one significant figure is therefore being applied to quantities that outsiders can rarely pin down to better than a factor of two.

For that reason, a compute threshold can serve as an administrable screening rule without becoming a scientific definition of frontier. Good policy design can combine input proxies with capability evaluations, deployment context, reach, safeguards, and a mechanism for updating or rebutting the classification.

The frontier in mid-2026

Any roster of leading models is a dated snapshot, and this one should be read as a reading of two public instruments on 1 August 2026 rather than as a ranking of laboratories. Independent aggregate measurement is preferable to vendor claims here, and the two most widely cited aggregators disagree, which is itself informative.

Artificial Analysis publishes an Intelligence Index that combines nine evaluations; version 4.1 of the index was in use, with a cost-methodology update noted on 30 July 2026. Its leaderboard was headed by Claude Opus 5 in a maximum-effort configuration at 61, followed by Claude Opus 5 at other reasoning settings, Claude Fable 5 at 60, and GPT-5.6 Sol at 59. Kimi K3 placed seventh at 57 and Grok 4.5, listed under the creator name SpaceXAI, placed eleventh at 54. The highest-placed Google model at that reading was Gemini 3.5 Flash at 50, in nineteenth place, just ahead of Gemini 3.6 Flash in twentieth.[34]

Arena, the human pairwise-preference leaderboard formerly published at lmarena.ai, ordered the same period differently. Its text board was headed by claude-fable-5 at 1509 plus or minus 6, with Anthropic models holding seven of the top ten places, a Meta model in eighth place at 1490, and Gemini 3 Pro tenth at 1486 plus or minus 4.[35]

The divergence is not a contradiction. One instrument aggregates scored benchmarks under controlled harnesses; the other aggregates human preference votes on open-ended chats. They also cover different releases at different times, since neither evaluates every configuration of every model on release day. A claim that a particular model "is the frontier model" is therefore incomplete without naming the instrument and the date, and the same two readings would support at least two different answers.

Chinese laboratories are directly relevant to the article's central point that frontier is relational. Epoch AI reported on 2 January 2026 that Chinese models had lagged the US frontier by about 7 months on average since 2023, with a range of 4 to 14 months, and that every model at the capability frontier since 2023 had been developed in the United States.[30] Seven months later, the top open-weights entry on the Artificial Analysis index was Kimi K3 from Moonshot AI at 57, four points below the leading closed model, with GLM-5.2 at 51 and a DeepSeek V4 Flash release at 50.[34] Whether that constitutes closing the gap depends entirely on whether the question is about the single best system, about the best freely downloadable one, or about compute.

How frontier status is assessed

No single procedure answers every frontier question. A reproducible assessment starts by defining the decision the evidence will inform. Comparative research may ask which system performs best on a task. A release review may ask whether a system crosses a hazardous-capability threshold. A regulator may ask whether a legal presumption or designation criterion is met.

A minimal assessment record includes:

  1. Object tested. Record the model identifier, weights or service version, date, system prompt, sampling parameters, inference budget, tools, retrieval, memory, and access controls.
  2. Reference class. State which other systems were eligible for comparison and why. Closed systems unavailable to the evaluator create an evidence boundary, not proof of inferiority.
  3. Capability or risk construct. Define what the evaluation is intended to measure and how success relates to real-world behavior.
  4. Elicitation protocol. Record prompting, fine-tuning, scaffolding, human assistance, tool access, and the search budget used to elicit performance.
  5. Metric and uncertainty. Report sample size, uncertainty intervals where meaningful, failure severity, and sensitivity to reasonable protocol changes.
  6. Data integrity. Investigate contamination, memorization, test leakage, grader error, and whether the tasks remain discriminating.
  7. External validity. Explain how the controlled task differs from deployment, including user skill, environmental complexity, and opportunities for repeated attempts.
  8. Review lifecycle. Reassess after material model, system, or deployment changes and incorporate incidents and field evidence.

These records prevent several common category errors. A capability evaluation estimates whether a system can produce a result under specified conditions. A risk assessment also considers who can access that capability, how it could cause harm, how likely the pathway is, the scale of consequences, and whether safeguards interrupt it. A model can cross a capability threshold without the evaluator having established the probability of a particular harm.

The 2026 International AI Safety Report calls the mismatch between pre-deployment tests and real-world behavior an "evaluation gap." It identifies narrow tasks, outdated tests, data contamination, inconsistent outputs, and differences between controlled and operational settings as recurring limitations.[6] Model evaluation therefore supports evidence gathering; it does not remove the need for threat modeling, field monitoring, and judgment.

Capability evaluations and red teaming

Capability evaluations measure performance on predefined tasks, including potentially hazardous tasks. They may establish lower bounds: if a system succeeds under a documented protocol, it has demonstrated at least that capability. A failure is harder to interpret because the evaluator may not have found an effective prompt, scaffold, or fine-tuning method.

Red teaming takes a more adversarial approach. Testers search for harmful behavior, policy bypasses, security weaknesses, or unexpected interactions. External red teams can add expertise and reduce some internal blind spots, but access restrictions, short testing windows, and incomplete knowledge of training and system controls can limit what they find.

Evaluation should be repeated at more than one stage. Testing before training can inform threat models and data controls. Checkpoints can reveal capability trends. Pre-release testing can inform deployment decisions. Post-deployment monitoring can reveal misuse, failures, and user adaptations absent from laboratory tests. Independent replication is valuable when it is technically and legally possible.

Risk and governance rationale

Frontier governance concentrates attention where a broadly capable model or system might enable unusually severe harm, where the developer has information unavailable to outsiders, or where failures could propagate widely. This focus does not imply that smaller or older systems are safe, nor that frontier risks are the only important AI harms. Bias, privacy loss, unsafe automation, labor effects, fraud, and environmental costs can arise far below any frontier threshold.

Risk discussions commonly separate several pathways:

  • Malicious use. A user may apply a capable system to cyber operations, chemical or biological work, fraud, coercion, or influence operations.
  • Loss of control or unintended behavior. A system may pursue an objective in an unsafe way, exploit an evaluation, circumvent oversight, or behave differently after deployment.
  • Model theft and proliferation. Unauthorized access to weights or development infrastructure can transfer capability and defeat a provider's access controls.
  • Deployment and ecosystem effects. Broad use can create correlated failures, concentration, labor-market disruption, information pollution, or dependencies in critical services.

Evidence varies substantially by pathway. The 2026 international report found robust empirical evidence for some present harms, while evidence for other severe scenarios relies more heavily on controlled studies, modeling, and theory. It also concluded that no current safeguard combination is perfectly reliable and described layered controls as a defense-in-depth strategy.[6] Responsible writing preserves that uncertainty instead of converting a plausible pathway into a forecast.

The phrase systemic risk is itself context-dependent. The EU AI Act defines it in Article 3(65) as a risk specific to the high-impact capabilities of general-purpose AI models that has a significant effect on the Union market because of their reach, or through reasonably foreseeable negative effects on public health, safety, public security, fundamental rights, or society as a whole, and that can propagate at scale across the value chain.[8] The International AI Safety Report uses systemic risks more broadly for effects arising from widespread deployment across society and the economy, and explicitly notes the difference.[6] The two senses should not be silently merged.

AI safety and AI governance also operate at different layers. Technical work can improve evaluation, robustness, interpretability, security, or monitoring. Organizational controls assign decision rights, escalation paths, documentation duties, and resources. Public governance can create reporting, audit, liability, standards, or enforcement mechanisms. A frontier designation by itself supplies none of these controls.

The instruments below are described as they stood on 1 August 2026. They differ in legal force, in what triggers coverage, and in what follows from coverage. Thresholds, guidance, and effective dates change, so the cited primary text controls.

InstrumentLegal statusTriggerCore consequence
Bletchley Declaration, November 2023Political declarationHighly capable general-purpose models and relevant narrow AI that match or exceed then-leading capabilities and could cause harmCooperation, testing, evaluation, transparency; no licensing class or compute threshold[2]
Frontier AI Safety Commitments, Seoul, May 2024Voluntary, 20 signatory organisationsFrontier AI defined relationally, with thresholds set by each signatoryPublish a safety framework, define intolerable-risk thresholds, and refrain from development or deployment if mitigations cannot hold risk below them[18]
EU AI Act, Regulation (EU) 2024/1689Binding EU regulationGPAI model with systemic risk: high-impact capabilities, presumed above 10^25 FLOP cumulative training compute, or Commission designation on Annex XIII criteriaArticle 55 evaluation and adversarial testing, systemic-risk assessment and mitigation, serious-incident reporting, cybersecurity[8]
US Executive Order 14110, October 2023Revoked 20 January 2025Dual-use foundation models above 10^26 FLOP, or above 10^23 FLOP if trained primarily on biological sequence data; clusters above 10^20 FLOP per secondReporting to the federal government; no longer in force[11][12]
US Executive Order 14409, June 2026In force"Covered frontier model" designated by the Director of the NSA on a cyber-capability assessmentVoluntary framework for federal evaluation access; expressly no mandatory licensing or preclearance[22]
California SB 1047, 2024Vetoed 29 September 2024Covered model above 10^26 FLOP with training cost above 100 million dollarsWould have required safety protocols and a Board of Frontier Models; never took effect[19][20]
California SB 53, Chapter 138, 2025EnactedFrontier model above 10^26 FLOP; heightened duties for large frontier developers above 500 million dollars in annual gross revenuePublished framework, transparency reports, 15-day critical-incident reporting, whistleblower protection, civil penalties up to 1 million dollars per violation[13]
New York General Business Law Article 44-B (RAISE Act)Enacted, effective 1 January 2027Frontier model above 10^26 FLOP; large frontier developer above 500 million dollars in annual gross revenueTransparency, reporting, governance, and large-developer disclosure duties[14][15]

The European Union

The EU thresholds are often misstated. Article 51 does not declare every model above 10^25 FLOP a frontier model. Article 51(1)(a) classifies a general-purpose AI model as carrying systemic risk when it has high-impact capabilities evaluated with appropriate technical tools, and Article 51(2) creates a presumption that this test is met when cumulative training compute exceeds 10^25 floating-point operations. Article 51(1)(b) lets the Commission classify a model on its own initiative, or after a qualified alert from the scientific panel, using the Annex XIII criteria; those criteria are the number of parameters, dataset size or quality, training compute or proxies for it such as cost, time, or energy, input and output modalities with modality-specific state-of-the-art thresholds, benchmark and evaluation results including autonomy and tool access, high market impact presumed at 10,000 or more registered business users in the Union, and the number of registered end users. Article 51(3) empowers the Commission to amend the thresholds by delegated act as hardware and algorithms improve.[8]

Article 52 supplies the procedure. A provider whose model meets the Article 51(1)(a) condition must notify the Commission without delay and in any event within two weeks, and may argue at the same time that the model exceptionally does not present systemic risks because of its specific characteristics. If the Commission rejects that argument the model is treated as carrying systemic risk, and a provider designated ex officio may later request reassessment.[8] The threshold is therefore a rebuttable presumption inside a notification procedure, not a definition of a product category.

Article 55 attaches four obligations to providers of GPAI models with systemic risk: model evaluation under standardised state-of-the-art protocols including documented adversarial testing, assessment and mitigation of Union-level systemic risks, tracking and reporting of serious incidents to the AI Office without undue delay, and an adequate level of cybersecurity for both the model and its physical infrastructure.[8]

Separately, nonbinding Commission guidelines use training compute above 10^23 FLOP plus specified generative modalities as an indicative way to identify a general-purpose AI model at all. The guidance says models above the indicator can exceptionally fall outside the class and models below it can still qualify.[9] The 10^23 figure is therefore neither the Act's statutory definition of GPAI nor a lower frontier tier.

The AI Act entered into force on 1 August 2024, the prohibitions applied from 2 February 2025, and the GPAI obligations applied from 2 August 2025. The Commission's enforcement powers and the Article 50 transparency rules apply from 2 August 2026. The Digital Omnibus package moved the high-risk deadlines back, to 2 December 2027 for the standalone Annex III systems and 2 August 2028 for high-risk AI embedded in regulated products, but did not move the general-purpose AI dates.[39] These dates describe the EU regime and should not be generalized to other jurisdictions.

United States

Federal policy on this point has reversed twice. Executive Order 14110 of 30 October 2023 introduced the term dual-use foundation model and set interim reporting triggers pending technical conditions from the Department of Commerce: training runs above 10^26 integer or floating-point operations, above 10^23 operations for models trained primarily on biological sequence data, and computing clusters co-located in a single data center with networking above 100 Gbit/s and theoretical capacity above 10^20 operations per second.[11] The order was revoked on 20 January 2025, and three days later Executive Order 14179, "Removing Barriers to American Leadership in Artificial Intelligence," set out the replacement policy without reinstating any reporting threshold.[12][41]

Two later orders shape the current position. Executive Order 14365, "Ensuring a National Policy Framework for Artificial Intelligence," signed 11 December 2025, directs the Attorney General to stand up an AI Litigation Task Force to challenge conflicting state AI laws, directs the Secretary of Commerce to identify onerous state laws, contemplates conditioning certain federal funds, and asks for legislation establishing a uniform federal framework.[21] Executive Order 14409, "Promoting Advanced Artificial Intelligence Innovation and Security," signed 2 June 2026, uses the phrase covered frontier model but delegates the classification rather than defining it numerically: the determination is made by the Director of the National Security Agency in consultation with the National Cyber Director, the Assistant to the President for Science and Technology, and the Director of CISA, and the participation route is a voluntary framework under which developers may seek designation and provide federal access to such models. The order states that nothing in the section authorises "a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models."[22]

The practical consequence is that the federal compute-reporting trigger created in 2023 has not been reinstated, and the only federal use of frontier model in force is a voluntary, security-led designation with no published numeric threshold. Writing that describes a US federal compute reporting duty as current is describing a revoked order.

State law

California SB 1047 would have created the first compute-plus-cost threshold in US law: a covered model was one trained with more than 10^26 integer or floating-point operations at a cost exceeding 100 million dollars, with a separate fine-tuning branch at 3x10^25 operations and 10 million dollars.[19] Governor Newsom vetoed it on 29 September 2024 with a message that is itself an argument about the concept. He wrote that "by focusing only on the most expensive and large-scale models, SB 1047 establishes a regulatory framework that could give the public a false sense of security about controlling this fast-moving technology," and that "smaller, specialized models may emerge as equally or even more dangerous than the models targeted by SB 1047."[20] The veto is the clearest official statement that a compute threshold selects for expense rather than for danger.

California SB 53, approved on 29 September 2025 as Chapter 138 and titled the Transparency in Frontier Artificial Intelligence Act, took the transparency half of that agenda without the liability half. It defines a frontier model as a foundation model trained using more than 10^26 integer or floating-point operations, and adds a second, firm-level filter: a large frontier developer is a frontier developer whose group had annual gross revenues above 500 million dollars in the preceding calendar year. Catastrophic risk is defined as a foreseeable and material risk of more than 50 deaths or serious injuries, or more than 1 billion dollars in property damage. Duties include publishing a frontier AI framework, publishing a transparency report at or before deployment, reporting critical safety incidents within 15 days of discovery, and whistleblower protections, with civil penalties up to 1 million dollars per violation. The Department of Technology is directed to recommend definitional updates annually from 1 January 2027.[13]

New York's RAISE Act, codified as General Business Law Article 44-B, is marked effective 1 January 2027. The text as revised on 3 April 2026 defines a frontier model as a foundation model trained using more than 10^26 operations, defines a large frontier developer by the same 500 million dollar revenue test, and uses the same 50-casualty or 1 billion dollar catastrophic-risk definition.[14][15]

The convergence is worth stating precisely. Two states now use the same order of magnitude and the same revenue filter, which reduces compliance divergence but does not make 10^26 FLOP a scientific finding. Each statute separately specifies which computation counts, which activities are covered, and what duties follow, and both pair the compute test with a revenue test, meaning the class is defined jointly by model scale and firm size rather than by capability. A well-funded developer training just below the line is outside the class; a small laboratory that somehow trained above it would be a frontier developer without being a large one.

International processes

The summit series that produced the vocabulary has continued and broadened. Bletchley Park hosted the first AI Safety Summit on 1 and 2 November 2023; Seoul followed in May 2024 and produced the Frontier AI Safety Commitments; Paris hosted the AI Action Summit in February 2025; and the India AI Impact Summit met at Bharat Mandapam in New Delhi in February 2026, publishing a set of outcome documents that includes the New Delhi Frontier AI Impact Commitments and separate Frontier AI Model Usage Commitments.[38] The four meetings are treated together as the AI Safety Summit series, though only the first carried that name.

National institutes have also changed their framing. The UK renamed its AI Safety Institute the AI Security Institute on 14 February 2025, narrowing its remit to risks with security implications such as chemical and biological weapons, cyber attacks, fraud, and child sexual abuse material, and stating that it would not focus on bias or freedom of speech.[37] The rename is a useful reminder that even the bodies created to evaluate frontier models disagree about which boundary matters.

The International AI Safety Report, mandated by the Bletchley summit, is the closest thing to a shared scientific baseline. The 2026 edition was published on 3 February 2026, chaired by Yoshua Bengio, with an Expert Advisory Panel nominated by 29 nations plus the UN, the OECD, and the EU, and contributions from more than 100 experts.[6]

Evaluation and assurance

Assurance is a body of evidence that a model or system satisfies stated claims under stated conditions. It is broader than passing a benchmark. For frontier systems, evidence may include development documentation, capability evaluations, adversarial testing, security assessments, incident processes, independent review, and post-deployment monitoring.

NIST's Generative AI Profile applies the functions Govern, Map, Measure, and Manage across the AI lifecycle. It identifies risks including confabulation, data privacy, harmful bias, information integrity, information security, intellectual property, and value-chain integration. The profile is voluntary and cross-sectoral; it is not a frontier classification or proof that a system is safe.[16] Its value here is procedural: risk management should connect measurement to context, responsibility, mitigation, and monitoring.

The EU AI Act makes some assurance practices legal obligations for providers of GPAI models with systemic risk under Article 55.[8] The Commission's General-Purpose AI Code of Practice, published on 10 July 2025, is a voluntary compliance tool in three chapters covering transparency, copyright, and safety and security; the Safety and Security chapter is aimed at providers subject to the Article 55 duties. Signing the code is not the source of the statutory obligation, and signature is chapter-selective in practice: the Commission's signatory list records xAI as having signed the Safety and Security chapter only.[10]

The Seoul Frontier AI Safety Commitments are a separate voluntary instrument, signed by 20 organisations. Participating organizations committed to publish safety frameworks focused on severe risks, assess risks across the lifecycle, define thresholds at which severe risks would be deemed intolerable, specify mitigations and escalation processes, assign accountability, and provide public or trusted-party transparency, and to refrain from developing or deploying a model at all if mitigations cannot keep risks below those thresholds. The commitments defined frontier AI relationally and allowed thresholds to use capability, risk, safeguards, deployment context, or other factors.[18]

Company frameworks

Those commitments produced a family of published documents. They are the operative definitions of frontier for the organizations that actually decide what gets trained and released, and the trigger is mostly, but not entirely, capability rather than compute. Anthropic's and Google DeepMind's frameworks are triggered purely by capability. Meta added a second, alternative trigger in April 2026: a model counts as Frontier if it shows high capabilities in catastrophic risk areas or if it was trained with at least 10^26 integer or floating-point operations.[26] OpenAI keeps compute-threshold logic in a separate document, the Frontier Governance Framework of 28 May 2026, and states there that the Preparedness Framework "does not depend on specific legal compute thresholds like the FGF." So the mismatch with statute is real but narrower than a clean capability-versus-compute split suggests: where a developer does adopt a compute number, it is the same 10^26 figure the statutes use. All four are revisable by their authors, and their risk domains and evidentiary standards differ.

FrameworkCurrent versionTrigger structureRisk domains named
Responsible Scaling Policy, Anthropic3.4, effective 8 July 2026Capability or usage thresholds mapped to mitigations, with a separate column of recommended industry-wide practice and competitor-contingent commitmentsNon-novel and novel chemical or biological weapons production; misaligned systems in high-stakes settings; AI research and development automation[23]
Preparedness Framework, OpenAI2, 15 April 2025High capability requires safeguards before deployment; Critical capability requires safeguards during development regardless of deployment plansTracked: biological and chemical, cybersecurity, AI self-improvement. Research: long-range autonomy, sandbagging, autonomous replication and adaptation, undermining safeguards, nuclear and radiological[24]
Frontier Safety Framework, Google DeepMind3.1, 17 April 2026Tracked Capability Levels and Critical Capability Levels, each with a recommended security level and alert thresholdsCBRN, cyber, harmful manipulation, machine learning research and development, misalignment[25]
Advanced AI Scaling Framework, Meta2, 7 April 2026, previously the Frontier AI Framework of 3 February 2025Outcomes-led thresholds of critical, high, and moderate or lower, keyed to whether the model would substantially contribute to a modelled threat scenarioCybersecurity, chemical and biological, loss of control[26]

Some details matter for reading them accurately. Anthropic's policy sets its AI research threshold as models that could fully substitute for the company's own research scientists and engineers at costs within a factor of five, or that produce a dramatic acceleration in the pace of AI progress, and it maps thresholds to protections at or above its ASL-3 standard rather than to a ladder of numbered levels; version 3.4 also states that the policy is not designed to satisfy statutory definitions and that obligations such as those under SB 53 are handled in separate compliance documents.[23] OpenAI defines severe harm as the death or grave injury of thousands of people or hundreds of billions of dollars of economic damage, and includes a marginal-risk clause allowing it to relax safeguards if a competitor releases a comparable high-capability system without comparable safeguards, subject to public acknowledgement and a commitment to remain more protective than that competitor.[24] Google DeepMind pairs each critical level with a recommended minimum security level for the field, on the reasoning that weight exfiltration is the dominant risk at some thresholds and not others.[25] Meta's document commits to continuing development past the critical threshold only after risk has been reduced to moderate or lower, with weight access controls overseen by named executives in the interim.[26]

Company-authored frameworks can turn evaluations into conditional development and deployment decisions, but they are not uniform standards. Their thresholds, evidence requirements, exception processes, and governance differ, and their authors can revise them; two of the four documents above were retitled or renumbered within eighteen months. Independent scrutiny therefore asks both whether a framework is well designed and whether it was followed in practice.

What assurance cannot establish

An evaluation result is bounded by its protocol. It generally cannot prove the absence of a capability, guarantee behavior under all prompts, or predict every downstream integration. Sparse observations are particularly weak evidence for rare events. Evaluator access may also omit training data, intermediate checkpoints, system prompts, monitoring data, or security architecture.

Useful assurance claims are consequently narrow and auditable. They state the model and configuration, the risk question, the evidence collected, known limitations, residual risk, the person or body making the decision, and the conditions that trigger reassessment. A model card can communicate part of this record, but publication format alone does not establish completeness or independent verification.

Open weights and downstream systems

Release mode is independent of frontier status. A model can be frontier and available only through a controlled service, frontier and released with downloadable weights, or open-weight but no longer near a selected frontier. "Open source" also has license and access meanings that should not be inferred merely from weight availability; see open weights.

The gap between the two release modes is measurable and has narrowed. Epoch AI estimated in October 2025 that open-weight models lagged the state of the art by about 3.5 months and 7 index points, and in May 2026 put the lag at about 4 months and 8 index points, with a 90 percent interval of 7 to 11 points.[29] On the Artificial Analysis index reading of 1 August 2026, the leading open-weights entry scored 57 against 61 for the leading closed model.[34] A lag of that size means the practical question is rarely whether an open-weight model is at the frontier, but how long the frontier stays exclusive.

The concrete cases differ in important ways. Meta's Llama family established large open-weight releases from a US laboratory as routine. DeepSeek, Alibaba's Qwen, Moonshot AI's Kimi, and Zhipu's GLM lines made the leading open-weight tier predominantly Chinese, which is why Epoch's US-China analysis and its open-versus-closed analysis measure overlapping but distinct things.[30][29] Licences vary and often are not standard open-source licences: the Kimi K3 repository, for example, distributes weights under a model-specific licence rather than a recognised open-source one.[40] OpenAI's gpt-oss release in August 2025 was accompanied by an unusual public artifact: a study in which the developers deliberately fine-tuned the model for maximum biological and cyber capability, in a reinforcement-learning environment with browsing for biology and an agentic capture-the-flag environment for cyber, and reported that the adversarially trained version still underperformed OpenAI o3, a model below the Preparedness Framework's High capability threshold, and did not substantially advance open-weight capability in either domain.[36]

That study illustrates what the frameworks now require. Meta's framework states that for open-weight releases and fine-tuning APIs it will model adversaries capable of modifying model behaviour through continued training, and that where it is considering releasing weights it will conduct domain-specific capability training to attempt to upper-bound the model's capabilities.[26] Under Google DeepMind's framework the recommended security level attaches to the capability level rather than to the release plan, on the reasoning that the value of weight security depends on what the weights can do.[25] The common structure is that release mode changes the elicitation standard rather than the threshold, though DeepMind's framework also allows a recommended security level to be relaxed where a model is not meaningfully more capable than other publicly available models, or where it judges the benefits of releasing the weights to outweigh the risks.

Open weights can support independent research, reproducibility, local adaptation, competition, privacy-preserving deployment, and access for less-resourced organizations. They also change the control surface. Once weights are widely distributed, the original provider generally cannot recall every copy, monitor every use, or enforce service-level safeguards. Users can fine-tune the model, remove filters, combine it with tools, or deploy it in environments unknown to the original developer. The net risk depends on the model's capabilities, the marginal assistance it provides over other resources, the surrounding ecosystem, and available defenses.[6]

Law treats release modes inconsistently. The EU AI Act exempts providers of qualifying free and open-source GPAI models from the technical-documentation duties in Article 53(1)(a) and (b), but the same paragraph states that the exception does not apply to general-purpose AI models with systemic risk, so an open-weight model above the Article 51 threshold carries the full Article 55 obligations.[8] The California and New York statutes take a different route: their definitions turn on training compute and developer revenue and say nothing about whether weights are published, so a sufficiently large open-weight release by a large developer is covered on the same terms as a closed one.[13][14] Neither approach reflects a finding that one release model is safer.

Downstream modification further complicates responsibility. Fine-tuning, retrieval, tool use, additional inference compute, or agent scaffolding can materially change behavior without retraining the original model from scratch. Risk controls may therefore attach to several actors: the original provider, a modifier, a system integrator, a deployer, an infrastructure provider, or a user. Which actor has a legal duty depends on the governing instrument; which actor can mitigate a risk depends on access to the relevant layer.

Interpretation and reporting

A frontier claim is most informative when it can be reconstructed. Reports should answer:

  • What exact model or system was assessed?
  • Which capability, risk, or legal definition was used?
  • What was the comparison set and cutoff date?
  • Was the claim based on training compute, observed capability, impact, reach, or a combination?
  • Which prompts, tools, safeguards, and inference resources were present?
  • How was uncertainty, contamination, and failed elicitation handled?
  • Is the result descriptive, voluntary, regulatory guidance, or legally binding?
  • What event would cause the classification to be reviewed?

Current-model rosters age poorly because releases, access conditions, and evaluation results change faster than a concept article can be reliably maintained. Epoch AI's finding that seventeen different models have held the top index position since March 2023, and that no run since GPT-4 has lasted even a third as long, is a quantitative statement of that problem.[32] Rosters can also imply a false total ordering across systems with different strengths, which is why the two aggregator readings above disagree about the same month. A dated, method-specific comparison belongs in a dedicated evaluation or timeline, with its inclusion rules and evidence exposed.

The most durable interpretation is therefore conditional: a frontier model is a model placed near a specified capability boundary or inside a specified policy or legal class. The definition, measurement conditions, and date are part of the claim, not optional footnotes.

References

  1. ^UK Department for Science, Innovation and Technology. "Frontier AI: capabilities and risks - discussion paper." Updated 28 April 2025. gov.uk/...-capabilities-and-risks-discussion-paper
  2. ^Countries attending the AI Safety Summit. "The Bletchley Declaration." 1 November 2023, updated 13 February 2025. gov.uk/...g-the-ai-safety-summit-1-2-november-2023
  3. ^Markus Anderljung et al. "Frontier AI Regulation: Managing Emerging Risks to Public Safety." 2023. arxiv.org/...2307.03718
  4. ^Rishi Bommasani et al. "On the Opportunities and Risks of Foundation Models." Stanford Center for Research on Foundation Models, 2021. crfm.stanford.edu/report
  5. ^OECD. "Explanatory memorandum on the updated OECD definition of an AI system." 5 March 2024. oecd.ai/...updated-oecd-definition-of-an-ai-system
  6. ^International AI Safety Report. "International AI Safety Report 2026." 3 February 2026. internationalaisafetyreport.org/...ety-report-2026
  7. ^Jaime Sevilla et al. "Compute Trends Across Three Eras of Machine Learning." 2022 International Joint Conference on Neural Networks, 2022. doi.org/...IJCNN55064.2022.9891914
  8. ^European Parliament and Council. Regulation (EU) 2024/1689, especially Articles 3(65), 51, 52, 53 and 55 and Annex XIII. 13 June 2024. eur-lex.europa.eu/...HTML
  9. ^European Commission. "Guidelines on obligations for General-Purpose AI providers." Updated 11 November 2025. digital-strategy.ec.europa.eu/...pose-ai-providers
  10. ^European Commission. "The General-Purpose AI Code of Practice." Updated 23 April 2026. digital-strategy.ec.europa.eu/...contents-code-gpai
  11. ^Executive Office of the President. Executive Order 14110, "Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence." 30 October 2023, published 1 November 2023. govinfo.gov/...2023-24283
  12. ^White House. "Initial Rescissions of Harmful Executive Orders and Actions." 20 January 2025. whitehouse.gov/...ful-executive-orders-and-actions
  13. ^California Legislature. Senate Bill 53, Chapter 138, "Artificial intelligence models: large developers" (Transparency in Frontier Artificial Intelligence Act). Approved 29 September 2025. leginfo.legislature.ca.gov/...billTextClient.xhtml
  14. ^New York State Senate. General Business Law Article 44-B, Section 1420, "Definitions." Revision dated 3 April 2026. nysenate.gov/...1420
  15. ^New York State Senate. General Business Law Article 44-B, "Responsible AI Safety and Education (RAISE) Act." Revision dated 3 April 2026. nysenate.gov/...A44-B
  16. ^National Institute of Standards and Technology. "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile." NIST AI 600-1, July 2024. doi.org/...NIST.AI.600-1
  17. ^Percy Liang et al. "Holistic Evaluation of Language Models." Transactions on Machine Learning Research, 2023. arxiv.org/...2211.09110
  18. ^UK Department for Science, Innovation and Technology and Republic of Korea. "Frontier AI Safety Commitments, AI Seoul Summit 2024." Updated 7 February 2025. gov.uk/...-safety-commitments-ai-seoul-summit-2024
  19. ^California Legislature. Senate Bill 1047 (2023-2024), enrolled text, "Safe and Secure Innovation for Frontier Artificial Intelligence Models Act." Enrolled 3 September 2024. leginfo.legislature.ca.gov/...billTextClient.xhtml
  20. ^Governor Gavin Newsom. Veto message returning Senate Bill 1047. 29 September 2024. gov.ca.gov/...SB-1047-Veto-Message.pdf
  21. ^Executive Office of the President. Executive Order 14365, "Ensuring a National Policy Framework for Artificial Intelligence." Signed 11 December 2025, published 16 December 2025. whitehouse.gov/...l-artificial-intelligence-policy
  22. ^Executive Office of the President. Executive Order 14409, "Promoting Advanced Artificial Intelligence Innovation and Security." Signed 2 June 2026, published 5 June 2026. whitehouse.gov/...lligence-innovation-and-security
  23. ^Anthropic. "Responsible Scaling Policy," version 3.4, effective 8 July 2026. anthropic.com/rsp
  24. ^OpenAI. "Preparedness Framework," version 2, 15 April 2025. cdn.openai.com/...preparedness-framework-v2.pdf
  25. ^Google DeepMind. "Frontier Safety Framework," version 3.1, published 17 April 2026. storage.googleapis.com/...safety-framework_3-1.pdf
  26. ^Meta. "Advanced AI Scaling Framework," version 2, 7 April 2026. ai.meta.com/...Meta_Advanced-AI-Scaling-Framework-v2
  27. ^Epoch AI. "Machine Learning Trends." Updated 5 February 2026. epoch.ai/trends
  28. ^Epoch AI. "Over 30 AI models have been trained at the scale of GPT-4." Updated 6 June 2025. epoch.ai/...models-over-1e25-flop
  29. ^Epoch AI. "Open models lag state-of-the-art closed models by 4 months," 29 May 2026, and "Open-weight models lag state-of-the-art by around 3 months on average," 30 October 2025. epoch.ai/...open-closed-eci-gap
  30. ^Epoch AI. "Chinese AI models have lagged the US frontier by 7 months on average since 2023." 2 January 2026. epoch.ai/...us-vs-china-eci
  31. ^Epoch AI. "The best score on the Epoch Capabilities Index grew almost twice as fast over the last two years as it did over the two years before that." epoch.ai/...ai-capabilities-progress-has-sped-up
  32. ^Epoch AI. "GPT-4 led in ECI far longer than any other model." 2 July 2026. epoch.ai/...gpt-4-longest-eci-lead
  33. ^Epoch AI. "The record for computing capacity in a single data center has doubled every 7 months." 11 June 2026. epoch.ai/...largest-data-center-compute
  34. ^Artificial Analysis. "Intelligence Index leaderboards," index version 4.1. Read 1 August 2026. artificialanalysis.ai/...models
  35. ^Arena (formerly LMArena). "Leaderboard." Read 1 August 2026. arena.ai/leaderboard
  36. ^Eric Wallace, Olivia Watkins, Miles Wang, Kai Chen and Chris Koch. "Estimating Worst-Case Frontier Risks of Open-Weight LLMs." arXiv, 5 August 2025. arxiv.org/...2508.03153
  37. ^UK Department for Science, Innovation and Technology. "Tackling AI security risks to unleash growth and deliver Plan for Change." 14 February 2025. gov.uk/...leash-growth-and-deliver-plan-for-change
  38. ^India AI Impact Summit 2026. "Outcome resources." February 2026. impact.indiaai.gov.in/outcome-resources
  39. ^European Commission. "AI Act." Regulatory framework page including the application timeline. digital-strategy.ec.europa.eu/...tory-framework-ai
  40. ^Hugging Face. "moonshotai/Kimi-K3" model repository. Read 1 August 2026. huggingface.co/...Kimi-K3
  41. ^Executive Office of the President. Executive Order 14179, "Removing Barriers to American Leadership in Artificial Intelligence." Signed 23 January 2025, published 31 January 2025. govinfo.gov/...2025-02172

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

9 revisions · v10 · 7,951 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independently fact-checked on 2026-08-01 against the Official Journal text of the EU AI Act, the executive orders on govinfo and the Federal Register API, the California and New York statutes, and the four developer safety frameworks. Every EU AI Act claim was confirmed, including that the 10^25 FLOP figure is a rebuttable presumption under Article 51(2) rather than a definition, and all 24 AI-related executive orders since January 2025 were enumerated to confirm that no federal compute-reporting threshold is currently in force. Six corrections were applied, the most significant being that the article stated no developer framework uses a compute threshold when Meta's Advanced AI Scaling Framework v2 added one in April 2026 and OpenAI maintains a separate compute-triggered document.

Cite this page: AI Wiki. "Frontier Models." aiwiki.ai, updated 1 Aug 2026, fact-checked 1 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/frontier_models

Suggest edit