Artificial General Intelligence

RawGraph

Artificial general intelligence (AGI) is a proposed form of artificial intelligence with broad, adaptable competence across many cognitive tasks, including tasks that were not anticipated during development. The term does not have a single scientific, legal, or industry-wide definition. Most definitions combine two ideas: breadth across domains and sufficiently high performance within those domains. They differ over the relevant human comparison group, the tasks that count, whether learning new tasks is essential, and whether physical action or autonomy is required.[1]

AGI is therefore a research goal and a contested category, not the name of a standardized test or an officially certified class of system. As of 28 July 2026, no generally accepted authority or measurement framework had established that an existing system met a broadly shared AGI threshold. Contemporary general-purpose AI systems can perform many well-scoped tasks at high levels, but their performance remains uneven across tasks and contexts. They still make basic errors, generate false statements, struggle with long projects, and perform less reliably outside controlled evaluations.[2]

The uncertainty is partly conceptual and partly empirical. Researchers disagree about what should count as general intelligence, how to measure it without rewarding memorization or benchmark-specific training, and which capabilities would be sufficient for reliable operation in the open world. Consequently, a claim that a system is "AGI" is only meaningful when it states a definition, a measurement procedure, the system conditions being tested, and the evidence supporting the claim.

Scope and terminology

Common elements

Although definitions vary, four recurring elements help distinguish AGI from a collection of unrelated task-specific skills:

  • Breadth: competence across a wide range of tasks or environments, rather than excellence in one narrowly specified task.
  • Depth: performance that reaches a stated threshold, often defined relative to humans with relevant skills.
  • Adaptation: the ability to learn, transfer knowledge, or acquire competence in tasks that were not individually represented in training.
  • Reliability: performance that remains useful under changes in wording, context, data distribution, and task duration.

Not every definition includes all four. OpenAI's charter, for example, defines AGI as "highly autonomous systems that outperform humans at most economically valuable work."[3] That is an organizational definition tied to work and autonomy. It is not a field-wide standard. A research framework from Google DeepMind instead separates performance and generality from autonomy, so that the capability level of a system can be discussed independently of how much freedom it is given when deployed.[1]

The word general also needs qualification. No finite system can be competent at every logically possible task. Human intelligence itself is shaped by sensory systems, culture, education, time, and physical constraints. Practical definitions therefore select a scope, such as cognitive tasks performed by skilled adults, economically valuable work, or performance across a specified distribution of environments.[4][5]

AGI is related to, but not interchangeable with, several other terms:

  • Narrow AI describes systems designed or optimized for a restricted task or domain. A narrow system may be far better than any human at its task without being general.
  • General-purpose AI describes models usable for many purposes. It is a broader product and policy category, and it does not by itself establish human-level adaptability or reliability. A large language model can be general-purpose while retaining highly uneven capabilities.
  • Artificial superintelligence usually means broadly capable intelligence that exceeds human performance, rather than merely matching a chosen human threshold. Some level-based schemes place it above AGI.[1]
  • Strong AI is a philosophical term often associated with the claim that a correctly programmed computer literally has a mind or understanding. John Searle's 1980 discussion of strong AI and the Chinese room concerned understanding and mentality, not a modern engineering threshold for broad task performance.[6]
  • Autonomy concerns how independently a system selects and carries out actions. A broadly capable model can be deployed with little autonomy, while a much narrower AI agent can be given substantial operational freedom.
  • Transformative AI refers to consequences, such as changes comparable in scale to major historical transformations. Such consequences might result from systems that do not satisfy a particular AGI definition.

AGI does not logically require consciousness, sentience, emotions, a human-like internal process, or a humanoid body. Some proposed definitions and research programs include one or more of these properties, but capability-based definitions generally do not.[1] Whether a machine could be conscious is a separate philosophical and scientific question.

History

The ambition behind AGI predates the term. In the 1955 proposal for the Dartmouth Summer Research Project on Artificial Intelligence, John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon proposed a 1956 study based on the conjecture that aspects of learning and intelligence could be described precisely enough for a machine to simulate them. The proposal listed language use, abstraction, problem solving, and self-improvement among its topics.[7]

In 1950, Alan Turing reframed the question "Can machines think?" as an observable imitation game. The later label "Turing test" refers to a conversational behavioral test derived from that proposal.[8] The test was historically influential, but conversational imitation does not by itself measure every capability associated with AGI. A system might imitate a human in dialogue without robust planning or physical competence, while a nonconversational system might have important forms of intelligence.

According to a historical review in the 2024 Levels of AGI paper, the phrase artificial general intelligence appeared in a 1997 paper by Mark Gubrud. The term became more organized as a research label in the following decade.[1] Ben Goertzel and Cassio Pennachin edited the 2007 volume Artificial General Intelligence, which collected work explicitly framed around broad machine intelligence.[9] The first AGI conference was held in Memphis, Tennessee, on 1-3 March 2008 in cooperation with the Association for the Advancement of Artificial Intelligence.[10]

The label gained wider use as machine learning systems became highly successful in specialized tasks. It offered a way to distinguish the older ambition of broad, adaptable intelligence from systems evaluated within a fixed domain. The boundary has remained disputed because systems can combine broad interfaces with narrow or brittle underlying competencies.

Definitions and operational frameworks

There is no definition-neutral AGI score. Every proposed measure embeds choices about tasks, prior knowledge, experience, human comparison groups, and acceptable error. Four influential approaches illustrate the range of those choices.

ApproachCore ideaWhat it contributesMain limitation
Performance and generality levelsClassify systems by both breadth and depth of capabilityReplaces a binary label with a progression and separates capability from autonomyRequires a broad, ecologically valid benchmark that does not yet exist
Universal intelligenceMeasure expected reward across a weighted distribution of computable environmentsProvides a mathematically explicit, agent-centered definitionThe ideal measure is not computable and cannot be implemented as a complete practical test
Skill-acquisition efficiencyMeasure how efficiently a system turns priors and experience into skill on new tasksDistinguishes intelligence from skill purchased with large amounts of data or task-specific trainingResults depend on how scope, priors, experience, and task difficulty are specified
Adult cognitive versatility and proficiencyCompare a system with well-educated adults across psychometrically motivated cognitive domainsMakes the human comparison and tested domains explicitIt is a proposed framework, not a consensus standard, and results depend on the selected tests and scoring rules

Levels of AGI

Morris and colleagues proposed a matrix with performance on one axis and generality on the other. Their performance levels run from no AI through emerging, competent, expert, exceptional, and superhuman. "Competent" means at least the 50th percentile of skilled adults, while "expert" and "exceptional" use the 90th and 99th percentiles. General systems must perform across a wide range of nonphysical tasks, including metacognitive tasks such as learning new skills.[1]

In the authors' September 2023 assessment, frontier language models were examples of "Emerging AGI," meaning broad but uneven performance comparable to or somewhat better than an unskilled human. They wrote that no public system had reached their "Competent AGI" cell at that time. These labels are properties of that framework, not declarations accepted by the field. The paper also treats autonomy as a deployment dimension rather than part of the capability score.[1]

Universal intelligence

Shane Legg and Marcus Hutter formalized intelligence as an agent's expected performance across environments, with simpler computable environments given greater weight. The proposal captures adaptation and goal achievement in a mathematically general way.[4] It also exposes a difference between a definition and a test: the ideal quantity depends on uncomputable Kolmogorov complexity, so any practical evaluation can only approximate it using a finite set of environments.

This approach does not select a uniquely correct human threshold for AGI. It instead supplies a general theory of machine intelligence. Translating that theory into an operational AGI decision still requires choices about environments, rewards, resources, and comparison groups.

Skill-acquisition efficiency

Francois Chollet argued that performance on a familiar task is a measure of skill, not necessarily intelligence. A system can obtain high task performance from extensive prior knowledge, large training sets, or task-specific engineering. His proposed measure defines intelligence in terms of the efficiency with which a learning system converts its priors and experience into skill on previously unknown tasks.[5]

This framing motivates evaluations that control what test-takers know before the test and how much experience they receive. It also makes comparisons conditional: two systems can be meaningfully compared only within a shared scope and with sufficiently comparable priors and experience. The ARC-AGI benchmark series grew out of this account of intelligence.

Cognitive-domain frameworks

A 2025 preprint by Dan Hendrycks and coauthors proposed defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. Its framework adapts ideas from Cattell-Horn-Carroll psychometrics and separates ten domains, including knowledge, reading and writing, mathematics, reasoning, working memory, long-term memory, visual processing, auditory processing, and processing speed.[11]

The paper's value is methodological: it states its human reference group and divides an aggregate label into inspectable components. Its reported model scores should not be treated as an official AGI percentage. The tests, weights, population baseline, model access conditions, and interpretation remain contestable, and the paper had not established a universal standard by July 2026.

Measuring progress

What a credible evaluation would need

A useful AGI evaluation would need evidence about more than peak accuracy on a collection of questions. At minimum, it would have to address:

  1. Task coverage: Does the test sample a defensible range of cognitive activities, rather than a convenient set of academic benchmarks?
  2. Novelty: Are tasks genuinely new to the system, or could their solutions, close variants, or grading patterns have appeared in training?
  3. Adaptation: Can the system learn new rules from limited experience and transfer what it learned?
  4. Robustness: Does performance persist across paraphrases, changed interfaces, unfamiliar contexts, and adversarial or accidental perturbations?
  5. Calibration and reliability: Does the system recognize uncertainty, avoid confident fabrication, and recover from errors?
  6. Resource accounting: How much data, computation, human scaffolding, tool access, time, and repeated sampling were used?
  7. Human comparison: Which humans are the baseline, what assistance do they receive, and how are differences in speed, cost, and background handled?
  8. External validity: Do controlled results predict performance in real tasks with ambiguous goals, interruptions, and consequences?

These requirements can conflict. A test broad enough to sample many activities may be difficult to secure against data contamination. A private test reduces leakage but makes independent replication harder. Human baselines improve interpretability but vary by education, culture, accessibility, and expertise. A single aggregate number can conceal severe weakness in a safety-critical domain.

Benchmark evolution

The ARC series illustrates why evaluation changes as systems and training methods change. The original ARC used novel grid-transformation tasks and a deliberately restricted set of assumed priors. ARC-AGI-2 increased the reasoning depth of the static tasks. ARC-AGI-3, introduced in 2026, moved to interactive, turn-based environments in which an agent must explore, infer goals, build an internal model, and plan without natural-language instructions.[12]

In the benchmark authors' March 2026 testing, humans solved all of the calibrated environments while the tested frontier systems scored below 1 percent. That is evidence about those systems under the benchmark's stated conditions, not proof that humans possess a single measurable essence of general intelligence or that a particular percentage would certify AGI. The paper itself emphasizes that static benchmarks can become vulnerable to direct optimization and overfitting as training data expands.[12]

Common interpretation errors

Several recurring errors make AGI claims appear stronger than their evidence:

  • Benchmark saturation: High performance can reflect training targeted at a known test rather than broad transfer.
  • Data contamination: Test items or close analogues may be present in training data.
  • Best-of-many reporting: Selecting the best sample or prompt hides the reliability experienced by a user who gets one attempt.
  • Tool and scaffold ambiguity: A result may depend on retrieval systems, code execution, human-written plans, or repeated retries that are not included in the model description.
  • Category substitution: Excellence in mathematics, coding, conversation, or game play is treated as proof of general competence.
  • Anthropomorphic inference: Fluent language is taken as direct evidence of understanding, stable beliefs, consciousness, or independent goals.
  • Deployment inference: Capability under laboratory conditions is assumed to imply economical, lawful, and dependable replacement of human work.

The 2026 International AI Safety Report calls the last group of problems an "evaluation gap": performance in pre-deployment tests does not reliably predict utility or risk in real-world settings. It identifies narrow task coverage, outdated evaluations, data contamination, prompt sensitivity, and controlled laboratory conditions among the causes.[2]

Current state

The most defensible description of current systems is capability-specific. General-purpose models can converse in many languages, generate software for bounded tasks, create media, and solve some advanced mathematics and science problems. Agents built around them can complete some useful computer tasks with limited oversight. Post-training methods and additional computation at inference time have produced substantial gains in reasoning-related evaluations.[2]

At the same time, capability profiles remain jagged. A system may solve a difficult formal problem and then fail on a simpler task because the wording, visual layout, required memory, or sequence length changes. Current systems still produce hallucinations, inconsistent answers, and mistakes that derail multi-step work. Performance declines on longer tasks, and reliable adaptation to unexpected obstacles remains limited. Integration with robotics also lags performance in text and other digital domains.[2]

For those reasons, the following statements are different:

  • A system can answer questions in many subjects.
  • A system can complete selected professional tasks under evaluation conditions.
  • A system can learn unfamiliar tasks efficiently.
  • A system can perform most cognitive work reliably and economically.
  • A system meets a particular published definition of AGI.
  • The field agrees that AGI has been achieved.

Evidence for one statement does not automatically establish the next. Claims about current AGI status should identify the exact system version, access mode, tools, evaluation dates, scoring procedure, and definition. Model branding or a developer's internal classification is not independent scientific validation.

Forecasts

Forecasts about AGI are statements of belief under uncertainty, not measurements of an existing capability. Their results depend strongly on definitions and question wording.

A survey reported by Katja Grace and colleagues collected responses from 2,778 researchers who had published in six major AI venues. Under a condition that scientific activity continued without major disruption, the aggregate forecast assigned a 10 percent probability by 2027 and a 50 percent probability by 2047 to "unaided machines" outperforming humans in every possible task. The median forecast for all human occupations becoming fully automatable was much later, reaching 50 percent in 2116.[13] Those two questions produced very different timelines even though they may appear related.

The authors reported substantial framing effects and explicitly cautioned that expert predictions should not be treated as objective truth. The event in their question is also not identical to every AGI definition. A survey can measure the distribution of respondents' beliefs, but it cannot establish the date at which a technological threshold will occur.

The 2025 AAAI Presidential Panel on the Future of AI Research reported a different kind of opinion evidence. Among 475 survey respondents, 76 percent judged that scaling current AI approaches to AGI was unlikely or very unlikely to succeed. Seventy-seven percent preferred designing systems with an acceptable risk-benefit profile over directly pursuing AGI, while 70 percent opposed halting AGI-directed research until full safety and control mechanisms were established.[14] These results document views in that survey sample. They do not prove either that scaling will fail or that AGI is feasible.

Public arrival dates stated by executives, investors, prediction markets, or individual researchers should be attributed to the speaker and dated. A moving forecast is not a fact about technical readiness. Articles about AGI should avoid presenting a consensus year when no such consensus exists.

Research approaches

No research program has been shown to be a necessary or sufficient route to AGI. Several partially overlapping approaches are active.

General models, post-training, and tools

One approach builds increasingly capable foundation models from broad data, then improves them through instruction tuning, preference optimization, reinforcement learning, synthetic data, tool use, retrieval, and additional computation during inference. This route has produced systems that reuse a shared model across language, coding, vision, audio, and other tasks. The 2026 International AI Safety Report attributes important recent capability gains to post-training and inference-time techniques, not only to larger initial training runs.[2]

Breadth of interface is not the same as robust generality. Research challenges include factual reliability, efficient learning from small amounts of experience, persistent memory, planning over long horizons, uncertainty estimation, and transfer under distribution shift. External tools can compensate for some weaknesses, but an evaluation must then specify the entire system, not only the base model.

Reinforcement-learning agents

Reinforcement learning studies agents that act in an environment to maximize a reward signal. David Silver, Satinder Singh, Doina Precup, and Richard Sutton advanced the position that reward maximization could, in principle, produce abilities associated with intelligence, including learning, perception, generalization, language, and social intelligence.[15] The paper presents a hypothesis, not an empirical demonstration that reward alone has yielded AGI.

Agent research makes goals, actions, feedback, exploration, and long-term consequences explicit. Its open problems include specifying rewards that reflect what people intend, avoiding shortcuts that score well without accomplishing the real task, learning efficiently in complex environments, and operating safely when feedback is sparse or delayed.

Cognitive architectures

Cognitive-architecture research tries to integrate memory, perception, action selection, learning, and reasoning in a coherent system. The Common Model of Cognition proposed by John Laird, Christian Lebiere, and Paul Rosenbloom identifies common structures across several human-like cognitive architectures, including working memory, declarative and procedural long-term memory, perception, and motor systems.[16]

Such architectures offer explicit hypotheses about how capabilities interact and how knowledge persists over time. Their components may be symbolic, neural, or hybrid. They do not by themselves establish that reproducing a selected cognitive structure will yield the breadth, efficiency, or reliability expected of AGI.

World models and embodiment

Another line of work emphasizes systems that learn predictive representations of their environments and use them for planning. Anna Dawid and Yann LeCun described hierarchical joint-embedding predictive architectures and latent-variable energy-based models as components of a proposed route toward autonomous machine intelligence.[17] This is one research proposal among several.

Embodied AI adds perception and action in physical or simulated environments. Embodiment can expose a system to causality, object persistence, space, and the consequences of action in ways that static text prediction does not. However, the claim that a physical body is necessary for AGI remains disputed. Capability frameworks can treat physical skill as an additional source of generality without making it a prerequisite for all cognitive definitions.[1]

Neurosymbolic and hybrid systems

Deep learning is effective at learning representations from large datasets, while symbolic AI can make rules, constraints, and structured knowledge explicit. Neurosymbolic research explores principled combinations of neural learning with symbolic knowledge representation and logical reasoning.[18] The goals include better systematic generalization, interpretability, and reasoning, but there is no agreed hybrid architecture that resolves all of those problems.

Related proposals combine learned perception with search, planning, causal models, program synthesis, external memory, or verification. The diversity of approaches reflects uncertainty about whether current model families need incremental improvements, new components, or fundamentally different learning principles.

Human-like learning

Some researchers argue that progress requires systems that learn more like people: from fewer examples, with compositional concepts, intuitive theories of objects and agents, and the ability to build causal explanations. A 2017 review by Brenden Lake and colleagues framed these properties as ingredients for machines that learn and think like people.[19] Human cognition is a source of hypotheses, not a specification that must be copied in every detail. Biological limits and biases are not automatically desirable engineering features.

Critiques of the goal

The usefulness of AGI as a target is itself disputed. A 2026 paper by Judah Goldfeder, Philippe Wyder, Yann LeCun, and Ravid Shwartz-Ziv argues that humans are not general in an unrestricted sense and proposes "superhuman adaptable intelligence" as an alternative goal centered on learning important skills and filling gaps in human capability.[20] This is a recent position paper, not a new consensus term.

Critics also argue that the binary AGI label encourages threshold marketing, hides uneven capability profiles, and centers human replacement rather than useful complementarity. Defenders reply that a shared term is useful for discussing systems whose breadth and adaptability could create qualitatively different opportunities and risks. Level-based and multidimensional frameworks are attempts to retain the topic while making claims more precise.

Safety and control

AI safety concerns current systems as well as possible future AGI. It is a mistake to postpone safety analysis until a disputed AGI threshold is crossed. Current general-purpose systems already create risks through misuse, unreliable outputs, privacy failures, bias, cyber operations, and effects on information environments. More capable systems may increase some risks or create new combinations of them.[2]

Early technical work on accident risks identified problems such as unintended side effects, reward hacking, inadequate oversight, unsafe exploration, and failure under distribution shift.[21] These are not unique to AGI. They become more consequential when a system has broader access, greater autonomy, or influence over high-stakes processes.

Alignment

AI alignment studies how to make system behavior accord with human intentions, constraints, and values. The problem includes at least three layers:

  • Specification: translating goals and rules into training signals or system requirements.
  • Generalization: maintaining intended behavior in situations not represented during development.
  • Control and oversight: detecting problems, limiting access, correcting behavior, and shutting systems down when necessary.

Human preferences are incomplete, context-dependent, and sometimes conflicting. Alignment is therefore not simply a matter of entering a perfect objective. It also involves institutional choices about whose interests count, how conflicts are resolved, and who can audit or contest a system.

Loss-of-control scenarios

The 2026 International AI Safety Report defines loss-of-control scenarios as cases in which one or more general-purpose AI systems operate outside anyone's control and regaining control is extremely costly or impossible. It reports wide expert disagreement about the likelihood of such scenarios.[2]

The report also draws an important boundary around present evidence. Current systems show early signs of some relevant capabilities in laboratory settings, but not at levels that would enable loss of control. Severe scenarios would require a combination of capabilities such as evading oversight, executing long-term plans, maintaining operation, and resisting countermeasures. Current agents remain unreliable on long tasks, and the integration and robustness required for those scenarios is beyond current systems.[2]

This does not establish that future loss of control is impossible. It means the evidence supports neither treating catastrophe as a demonstrated consequence of AGI nor dismissing it as logically excluded. Risk analysis must state assumptions about future capability, system behavior, access, deployment, and countermeasures.

Economic and social implications

AGI is often discussed as an economic threshold, but capability and labor substitution are not equivalent. A system might perform a task in a benchmark yet be too unreliable, expensive, slow, insecure, or difficult to integrate for production use. Conversely, a set of narrower systems may automate substantial work without satisfying an AGI definition.

The economic effect of advanced AI depends on the tasks systems can perform, adoption costs, organizational redesign, worker responses, regulation, market structure, and the distribution of gains. The 2026 International AI Safety Report describes current evidence on labor effects as mixed and the longer-term impact as uncertain.[2] This uncertainty does not justify attaching precise gross domestic product, unemployment, or company-valuation figures to AGI without a transparent model and a credible source.

Potential benefits include assistance in science, education, accessibility, medicine, engineering, and public services. Potential harms include displacement, wage and bargaining-power changes, concentration of economic and political power, unequal access, surveillance, and dependence on systems that are difficult to audit. None follows from the word AGI alone. Each depends on technical design and deployment institutions.

Governance and risk management

Because AGI lacks an agreed threshold, governance can use measurable properties instead of waiting for a label. Relevant properties include capability in hazardous domains, autonomy, access to tools and infrastructure, ability to copy or modify software, scale of deployment, and the severity of failure.

Risk-management measures can include pre-deployment evaluation, independent testing, access controls, staged release, monitoring, incident reporting, red-team exercises, cybersecurity, human authorization for consequential actions, and procedures for rollback or shutdown. The appropriate combination depends on the system and context. The International AI Safety Report describes defense in depth as layering safeguards because no single measure is perfectly reliable.[2]

The U.S. National Institute of Standards and Technology's AI Risk Management Framework is a voluntary framework for incorporating trustworthiness considerations into the design, development, use, and evaluation of AI systems. It was released in 2023, followed by a generative-AI profile in 2024; NIST states that AI RMF 1.0 is being revised.[22] The framework applies to AI risk management generally and does not certify AGI.

International coordination is difficult because development, deployment, and effects cross borders, while legal systems and social priorities differ. Information asymmetry also matters: developers may possess evidence about training, capabilities, incidents, and safeguards that external evaluators cannot access. Governance proposals therefore differ over reporting duties, audit access, compute oversight, liability, model release, and public participation.

Open questions

Key unresolved questions include:

  • What scope of tasks is broad enough to justify the word general?
  • Should AGI be measured relative to an average adult, skilled adults, expert groups, or the combined range of human institutions?
  • How should an evaluation account for tools, memory systems, repeated attempts, computation, energy, and human assistance?
  • Is fast learning on unfamiliar tasks more central than performance after extensive training?
  • How should physical skills and robotics affect a generality judgment?
  • Can a finite benchmark resist contamination and direct optimization for long enough to remain informative?
  • How should reliability, uncertainty calibration, and error recovery be included in a capability threshold?
  • Does autonomy belong in the definition, or should it remain a separate deployment property?
  • Which safety requirements should depend on measured capabilities rather than the AGI label?
  • Is a multidimensional profile more useful than a binary achieved/not-achieved judgment?

These questions are not peripheral details. They determine what an AGI claim means and what evidence could confirm or disconfirm it.

See also

References

  1. ^Meredith Ringel Morris et al., "Position: Levels of AGI for Operationalizing Progress on the Path to AGI," Proceedings of Machine Learning Research 235, 36308-36321 (2024).
  2. ^International AI Safety Report, International AI Safety Report 2026 (February 2026).
  3. ^OpenAI, "OpenAI Charter" (2018).
  4. ^Shane Legg and Marcus Hutter, "Universal Intelligence: A Definition of Machine Intelligence," Minds and Machines 17, 391-444 (2007).
  5. ^Francois Chollet, "On the Measure of Intelligence" (2019).
  6. ^John R. Searle, "Minds, Brains, and Programs," Behavioral and Brain Sciences 3(3), 417-457 (1980).
  7. ^John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon, "A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence" (1955).
  8. ^Alan M. Turing, "Computing Machinery and Intelligence," Mind 59(236), 433-460 (1950).
  9. ^Ben Goertzel and Cassio Pennachin, eds., Artificial General Intelligence (Springer, 2007).
  10. ^AGI-08: First Conference on Artificial General Intelligence (2008).
  11. ^Dan Hendrycks et al., "A Definition of AGI" (2025).
  12. ^ARC Prize Foundation, "ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence" (2026).
  13. ^Katja Grace et al., "Thousands of AI Authors on the Future of AI," Journal of Artificial Intelligence Research 84 (2025).
  14. ^AAAI Presidential Panel on the Future of AI Research, The Future of AI Research (2025).
  15. ^David Silver, Satinder Singh, Doina Precup, and Richard S. Sutton, "Reward is Enough," Artificial Intelligence 299, 103535 (2021).
  16. ^John E. Laird, Christian Lebiere, and Paul S. Rosenbloom, "A Standard Model of the Mind," AI Magazine 38(4), 13-26 (2017).
  17. ^Anna Dawid and Yann LeCun, "Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence" (2023).
  18. ^Artur d'Avila Garcez and Luis C. Lamb, "Neurosymbolic AI: The 3rd Wave" (2020).
  19. ^Brenden M. Lake et al., "Building Machines That Learn and Think Like People," Behavioral and Brain Sciences 40, e253 (2017).
  20. ^Judah Goldfeder, Philippe Wyder, Yann LeCun, and Ravid Shwartz-Ziv, "AI Must Embrace Specialization via Superhuman Adaptable Intelligence" (2026).
  21. ^Dario Amodei et al., "Concrete Problems in AI Safety" (2016).
  22. ^National Institute of Standards and Technology, "AI Risk Management Framework".

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

9 revisions · v10 · 4,953 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent 2026-07-28 fact-check: 22 primary, official, and peer-reviewed sources; definitions, dated forecasts, current capability, and safety boundaries verified.

Cite this page: AI Wiki. "Artificial General Intelligence." aiwiki.ai, updated 29 Jul 2026, fact-checked 29 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/artificial_general_intelligence

Suggest edit