History of Artificial Intelligence

RawGraph

Last edited

Fact-checked

Sources

37 citations

Revision

v1 · 3,687 words

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

The history of artificial intelligence is the story of a scientific field that has cycled between extravagant promises, painful funding collapses, and genuine breakthroughs since the mid-twentieth century. The field has a precise birthday of sorts: a summer workshop at Dartmouth College in 1956 where the term "artificial intelligence" was adopted. Its intellectual roots run earlier, through formal logic, the first mathematical model of the neuron in 1943, and Alan Turing's 1950 proposal that the question "Can machines think?" be replaced with a practical imitation game [1][2].

Twice the field promised more than it could deliver and paid for it with an AI winter, first after the 1973 Lighthill report and again when the 1980s expert-systems industry collapsed. Twice it rebuilt itself on new foundations: statistical machine learning in the 1990s, then deep learning after 2012, when a neural network called AlexNet won the ImageNet competition by a wide margin. The transformer architecture of 2017 and the large language models built on it produced ChatGPT in November 2022, which reached an estimated 100 million monthly users within about two months and set off the largest investment boom in the field's history [3][4]. As of 2026 that boom is still running, alongside an old and unresolved argument about how far the current techniques can go.

Precursors: logic, neurons, and Turing (1943-1955)

The idea that reasoning could be mechanized predates computers by centuries, but the immediate precursors of AI came out of wartime and postwar work on computation and the brain. In 1943 the neurophysiologist Warren McCulloch and the logician Walter Pitts published "A Logical Calculus of the Ideas Immanent in Nervous Activity," which treated networks of simplified all-or-none neurons as devices for computing propositional logic [1]. The paper supplied the founding abstraction of neural network research: intelligence, whatever else it is, might be the activity of many simple threshold units wired together. The paper circulated widely in the cybernetics community, which through the late 1940s and 1950s studied feedback and control in machines and living systems.

In 1950 Turing published "Computing Machinery and Intelligence" in the philosophy journal Mind. He argued that "Can machines think?" was too vague to answer and proposed instead the imitation game, later called the Turing test: a judge converses by teleprinter with a hidden computer and a hidden person, and the machine passes if the judge cannot reliably tell them apart. Turing predicted that by the end of the century machines would play the game well enough to make the question of machine thinking unremarkable [2].

The Dartmouth workshop (1956)

The field's founding document is a funding proposal dated August 31, 1955, written by John McCarthy, then a young Dartmouth mathematics professor, with Marvin Minsky, Nathaniel Rochester of IBM, and Claude Shannon of Bell Labs. It asked the Rockefeller Foundation to support "a 2 month, 10 man study of artificial intelligence" and stated the field's premise in one sentence: "The study is to proceed on the basis of the conjecture that every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it" [5]. McCarthy chose the name "artificial intelligence" partly to keep the new subject separate from cybernetics and automata theory [6].

The Dartmouth workshop ran for roughly eight weeks in the summer of 1956, from about June 18 to August 17 according to Ray Solomonoff's notes. Attendees came and went; the group included Ray Solomonoff, Oliver Selfridge, Arthur Samuel, Allen Newell, and Herbert Simon [6]. Newell and Simon arrived with something concrete: the Logic Theorist, written with J. C. Shaw, a program that proved theorems from Whitehead and Russell's Principia Mathematica and eventually handled 38 of the first 52 it attempted [3]. It is usually counted as the first working AI program.

Early optimism (1956-1973)

The two decades after Dartmouth produced a string of firsts, mostly funded by US defense agencies. Newell and Simon followed the Logic Theorist with the General Problem Solver, an attempt at a domain-independent reasoning engine based on means-ends analysis [3]. Samuel's checkers program, described in his 1959 paper "Some Studies in Machine Learning Using the Game of Checkers," improved with experience and reached respectable amateur play, an early demonstration of machine learning by name [3]. McCarthy invented Lisp, which became the field's standard language for decades.

Neural networks had their first boom in the same period. Frank Rosenblatt described the perceptron in a 1957 Cornell Aeronautical Laboratory report and a 1958 Psychological Review paper, and the US Office of Naval Research funded a hardware implementation, the Mark I Perceptron, first shown publicly in 1960 [7]. Press coverage was breathless: after a 1958 demonstration, The New York Times reported the Navy's expectation of a machine that "will be able to walk, talk, see, write, reproduce itself and be conscious of its existence" [7]. The enthusiasm ended abruptly in 1969 when Minsky and Seymour Papert published Perceptrons, a mathematical analysis of what single-layer networks cannot compute. Funding for connectionist research dried up for roughly a decade [3].

Two other systems from the era became cultural reference points. Joseph Weizenbaum's ELIZA, built at MIT between 1964 and 1966, imitated a Rogerian psychotherapist by reflecting a user's words back as questions. Weizenbaum was disturbed by how readily users treated the pattern-matching script as a sympathetic listener, a phenomenon now called the ELIZA effect, and he spent much of his later career as a critic of the field [8]. At Stanford Research Institute, Shakey the robot (1966-1972) combined perception, mapping, and the STRIPS planner to navigate rooms and push objects, the first mobile robot to reason about its own actions [3].

Not every early result held up. A 1966 report by the US Automatic Language Processing Advisory Committee (ALPAC) concluded, after roughly $20 million of government spending, that machine translation was slower, less accurate, and more expensive than human translation, and support for the field largely ended [9].

The Lighthill report and the first AI winter (1973-1980)

By the early 1970s the gap between predictions and results had become hard to ignore. In 1973 the British mathematician James Lighthill delivered "Artificial Intelligence: A General Survey" to the UK Science Research Council. The report was severe: research had failed to face the combinatorial explosion that makes toy-world methods useless on real problems, and the field's grander objectives were nowhere in sight [10]. The British government subsequently ended AI funding at most UK universities [10].

American funding fell at the same time for related reasons. The 1969 Mansfield Amendment pushed DARPA toward mission-oriented research and away from the undirected exploration that had sustained AI labs, and disappointing results from the Speech Understanding Research program at Carnegie Mellon contributed to deep cuts by 1974 [9]. The years from roughly 1974 to 1980 are now called the first AI winter: projects were cancelled, labs shrank, and "artificial intelligence" became a phrase researchers avoided in grant applications.

The expert systems boom and the second winter (1980-1993)

The field revived in the 1980s around a narrower and more commercial idea: instead of general intelligence, capture the rules a human specialist uses in one domain. Stanford's Heuristic Programming Project under Edward Feigenbaum had pioneered the approach in the 1960s and 1970s with DENDRAL, which identified organic molecules from mass spectrometry data, and MYCIN, Edward Shortliffe's system for diagnosing blood infections [11]. The commercial breakthrough was XCON (also called R1), deployed at Digital Equipment Corporation in 1980 to configure VAX computer orders; it reportedly saved the company $40 million a year [3].

An industry followed. By 1985 US corporations were spending over $1 billion a year on AI, much of it on in-house expert system groups and on specialized Lisp machines; two thirds of Fortune 500 companies applied the technology in daily business during the decade [9][11]. Japan's Fifth Generation Computer Systems project, launched in 1982 around logic programming and Prolog, prompted competitive government programs in the US and UK [9].

The bust arrived in 1987, when general-purpose workstations from Sun Microsystems and cheaper desktop computers made dedicated Lisp hardware pointless; a market worth about half a billion dollars was replaced within a year [9]. Expert systems themselves proved brittle and expensive to maintain, since every rule had to be hand-written and the systems could not learn. The Fifth Generation project wound down without meeting its goals, and DARPA's Strategic Computing Initiative cut AI spending sharply. The second AI winter stretched from the late 1980s well into the 1990s [9].

Ironically, the foundations of the next boom were laid during the bust. John Hopfield's work on associative memory networks (the Hopfield network) and Geoffrey Hinton's work with colleagues on the Boltzmann machine revived neural network research in the early 1980s, and in 1986 David Rumelhart, Hinton, and Ronald Williams popularized backpropagation, the algorithm that makes training multilayer networks practical [3][12]. Both Hopfield and Hinton would share the 2024 Nobel Prize in Physics for this line of work [12].

Quiet progress: statistical methods, Deep Blue, and Watson (1990s-2011)

Through the 1990s AI research continued under safer names: machine learning, pattern recognition, informatics. The field shifted from hand-coded knowledge toward statistical methods trained on data, with probabilistic reasoning and decision theory replacing brittle rule bases. Results accumulated with little fanfare. Gerald Tesauro's TD-Gammon, built at IBM in the early 1990s, taught itself backgammon through temporal-difference reinforcement learning and self-play, reaching just below the level of the best human players by 1993; its approach directly influenced later systems, including AlphaGo [13].

The decade's most public milestone was chess. IBM's Deep Blue lost a 1996 match to world champion Garry Kasparov 4-2, then won the May 1997 rematch in New York 3.5-2.5, the first defeat of a reigning world chess champion by a computer under tournament conditions [14]. Deep Blue was specialized search hardware rather than a learning system, but the symbolism registered far beyond the field.

IBM repeated the exercise in a messier domain in February 2011, when its Watson question-answering system beat Jeopardy! champions Ken Jennings and Brad Rutter over three broadcast days, finishing with $77,147 to claim the $1 million prize [15]. The follow-through was less happy: IBM's attempt to build a health business around Watson absorbed roughly $4 billion in development spending before the Watson Health assets were sold to the private equity firm Francisco Partners in 2022 for about $1 billion, an early case study in the gap between AI demonstrations and AI products [15].

The deep learning era (2012-2016)

The modern era began with a dataset. Fei-Fei Li started the ImageNet project in 2006 at Princeton, and the finished dataset held more than 14 million hand-labeled images; the associated ImageNet Large Scale Visual Recognition Challenge (ILSVRC) began in 2010 [16]. In 2012 AlexNet, a deep convolutional neural network built by Alex Krizhevsky, Ilya Sutskever, and Hinton at the University of Toronto and trained on two consumer Nvidia GTX 580 GPUs, won the challenge with a top-5 error rate of 15.3 percent, more than ten points better than the runner-up [17]. The margin convinced the computer vision community almost overnight, and by 2017 winning ILSVRC systems exceeded 95 percent top-5 accuracy and the benchmark was considered solved [16].

Deep learning then swept one field after another: speech recognition, machine translation, and natural language processing via word embeddings such as word2vec and recurrent architectures like the LSTM. Generative modeling advanced too, with Ian Goodfellow and colleagues introducing the generative adversarial network in 2014. Hinton, Yann LeCun, and Yoshua Bengio, who had kept neural network research alive through the lean years, shared the 2018 Turing Award for the work that made deep networks practical [18].

The era's defining spectacle came from DeepMind, the London lab co-founded by Demis Hassabis and acquired by Google in 2014. Its AlphaGo system combined deep policy and value networks with Monte Carlo tree search. In October 2015 it beat the European Go champion Fan Hui 5-0, the first time a program had beaten a professional on a full board without handicap [19]. In March 2016 it defeated Lee Sedol, one of the strongest players of his generation, 4-1 in Seoul; game 2 included an unconventional move that impressed professional observers, and Lee's game 4 win was the only game AlphaGo conceded in the match [20]. AlphaGo beat the world number one, Ke Jie, 3-0 in May 2017, and AlphaGo Zero, announced that October, learned entirely from self-play and beat the Lee Sedol version 100 games to none [19].

Transformers and large language models (2017-2022)

In June 2017 eight Google researchers posted "Attention Is All You Need", which introduced the transformer, an architecture built entirely on attention mechanisms with no recurrence or convolutions [21]. The design trained efficiently on parallel hardware and scaled better than anything before it; the paper became the common ancestor of essentially every subsequent frontier model, and the transformer the default architecture of the field.

OpenAI, founded in 2015, bet the architecture would keep improving with scale. GPT-1 (June 2018) established generative pre-training on unlabeled text; Google's BERT (October 2018) applied bidirectional transformer encoders to language understanding and was serving nearly all English Google Search queries by late 2020 [22][23]. GPT-2 (February 2019) was initially withheld in stages over misuse concerns, which itself made news, with the full 1.5 billion parameter model released that November [22]. GPT-3 (May 2020) jumped to 175 billion parameters and showed that a single model could perform new tasks from a few examples in its prompt, without task-specific training [24]. Empirical scaling laws relating model size, data, and compute to performance gave the strategy a quantitative footing, and the term foundation models was coined for the resulting general-purpose systems.

Two more pieces completed the recipe. InstructGPT (January 2022) showed that reinforcement learning from human feedback could turn a raw text predictor into a system that follows instructions [22]. And text-to-image generators such as DALL-E and Stable Diffusion put generative AI in front of a broad public in 2022. Deep learning was also producing hard scientific results: DeepMind's AlphaFold 2 effectively solved the protein structure prediction benchmark CASP in November 2020 with a median accuracy score of 92.4 GDT, and by July 2022 its public database held predicted structures for about 200 million proteins [25].

ChatGPT and the generative AI boom (2022-2024)

OpenAI released ChatGPT as a free research preview on November 30, 2022: a chat interface over a GPT-3.5 series model tuned with RLHF. It reached one million users in five days and an estimated 100 million monthly active users within about two months, making it the fastest-growing consumer internet application to that point [4]. The launch compressed a decade of gradual progress into a single product moment and triggered a global scramble.

The next eighteen months set the current competitive map. OpenAI released GPT-4 on March 14, 2023, a multimodal model whose parameter count it declined to disclose; Sam Altman put its training cost above $100 million [26]. Anthropic, founded by former OpenAI researchers including Dario Amodei, announced Claude the same day, trained with its constitutional AI method, though initially available only to approved users [27]. Meta's Llama weights, announced for researchers in February 2023, leaked online within a week; Meta then leaned into openness, releasing Llama 2 under a commercial license in July 2023 and seeding an open-weights ecosystem [28]. Google answered with Gemini in December 2023 [29]. OpenAI's Sora text-to-video model followed in 2024 [30].

Governments moved unusually fast by their own standards. The first global AI Safety Summit at Bletchley Park (November 1-2, 2023) produced a declaration on frontier AI risk signed by 28 countries and the EU, including both the United States and China [31]. The EU's Artificial Intelligence Act, the first comprehensive AI law, entered into force on August 1, 2024 with obligations phasing in through 2026 and beyond [32].

Late 2024 brought two markers of how far the field had come. OpenAI's o1-preview (September 2024) inaugurated reasoning models trained by reinforcement learning to produce a long chain of thought before answering; the full o1 model solved 83 percent of problems on the AIME mathematics competition versus 13 percent for GPT-4o [33]. And in October 2024 the Nobel committees recognized the field twice in one week: Hopfield and Hinton took the physics prize for foundational neural network discoveries, while Hassabis and John Jumper shared the chemistry prize with David Baker for protein structure prediction [12].

The frontier era (2025-2026)

The years since 2024 have been defined by three forces: reasoning models, open-weight challengers, and an infrastructure buildout with few precedents in industrial history.

In January 2025 the Chinese lab DeepSeek released DeepSeek-R1, an MIT-licensed open-weight reasoning model with performance comparable to leading Western systems, after claiming to have trained its V3 base model for about $5.6 million [34]. The release punctured assumptions about the capital required for frontier AI: on January 27, 2025, Nvidia lost roughly $600 billion in market value in the largest single-company one-day decline in US stock market history [34]. The scare proved temporary. Nvidia became the first company to reach a $4 trillion valuation on July 9, 2025 and passed $5 trillion on October 29, 2025 [30]. The Stargate venture, announced at the White House on January 21, 2025 by OpenAI, SoftBank, Oracle, and MGX, planned up to $500 billion of US AI data center construction by 2029 [35].

Model releases kept a punishing cadence. OpenAI's GPT-5 (August 7, 2025) unified fast and reasoning models behind an automatic router, to a mixed reception that fed a broader debate about diminishing returns from pre-training scale [36]. Google's Gemini 3 arrived on November 18, 2025, and Anthropic shipped Claude Sonnet 4.5 (September 2025) and Claude Opus 4.5 (November 2025), with the major labs increasingly marketing their models as agents that use computers, write software, and carry out long tasks rather than as chatbots [27][29]. The 2024 Turing Award went to Andrew Barto and Richard Sutton for reinforcement learning, the technique underneath both AlphaGo and the reasoning-model turn [18].

By 2026 ChatGPT alone was reported to serve around 900 million weekly users [4]. At the same time, an MIT study finding that 95 percent of business AI projects were unprofitable circulated widely, and arguments over an AI bubble, over artificial general intelligence timelines, and over AI safety continued in parallel with record capital spending [30]. The bulk of the EU AI Act's obligations are scheduled to apply from August 2026, and the follow-up India AI Impact Summit drew participants from more than 100 countries to New Delhi in February 2026 [32][37]. Seventy years after Dartmouth, the founding conjecture, that every feature of intelligence can be precisely enough described for a machine to simulate it, remains neither proved nor refuted; it is simply better funded than ever.

Timeline

YearEvent
1943McCulloch and Pitts publish the first mathematical model of neural computation [1]
1950Turing's "Computing Machinery and Intelligence" proposes the imitation game [2]
1956Dartmouth workshop; the field is named; Logic Theorist demonstrated [3][6]
1958Rosenblatt publishes the perceptron [7]
1966ELIZA published; ALPAC report ends most machine translation funding [8][9]
1969Minsky and Papert's Perceptrons stalls neural network research [3]
1973Lighthill report; first AI winter follows (roughly 1974-1980) [9][10]
1980XCON expert system deployed at DEC [3]
1982Japan launches the Fifth Generation Computer Systems project [9]
1986Rumelhart, Hinton, and Williams popularize backpropagation [3]
1987Lisp machine market collapses; second AI winter begins [9]
1997Deep Blue defeats Kasparov 3.5-2.5 [14]
2011IBM Watson wins Jeopardy! [15]
2012AlexNet wins ImageNet challenge; deep learning era begins [17]
2016AlphaGo defeats Lee Sedol 4-1 [20]
2017"Attention Is All You Need" introduces the transformer [21]
2020GPT-3 (175B parameters); AlphaFold 2 solves CASP [24][25]
2022ChatGPT launches November 30 [4]
2023GPT-4, Claude, Llama, Gemini; Bletchley Declaration [26][27][28][29][31]
2024EU AI Act in force; o1 reasoning model; Nobel Prizes for Hinton, Hopfield, Hassabis, Jumper [12][32][33]
2025DeepSeek-R1 and market shock; Stargate; GPT-5; Nvidia passes $5 trillion [30][34][35][36]
2026India AI Impact Summit in New Delhi (February); main EU AI Act obligations scheduled to apply from August [32][37]

See also

References

  1. McCulloch, Warren S. and Pitts, Walter. "A Logical Calculus of the Ideas Immanent in Nervous Activity." Bulletin of Mathematical Biophysics, Vol. 5, 1943 (reprint). https://www.cs.cmu.edu/~./epxing/Class/10715/reading/McCulloch.and.Pitts.pdf
  2. Wikipedia. "Computing Machinery and Intelligence." https://en.wikipedia.org/wiki/Computing_Machinery_and_Intelligence
  3. Wikipedia. "History of artificial intelligence." https://en.wikipedia.org/wiki/History_of_artificial_intelligence
  4. Wikipedia. "ChatGPT." https://en.wikipedia.org/wiki/ChatGPT
  5. McCarthy, J., Minsky, M., Rochester, N., Shannon, C. "A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence," August 31, 1955. https://www-formal.stanford.edu/jmc/history/dartmouth/dartmouth.html
  6. Wikipedia. "Dartmouth workshop." https://en.wikipedia.org/wiki/Dartmouth_workshop
  7. Wikipedia. "Perceptron." https://en.wikipedia.org/wiki/Perceptron
  8. Wikipedia. "ELIZA." https://en.wikipedia.org/wiki/ELIZA
  9. Wikipedia. "AI winter." https://en.wikipedia.org/wiki/AI_winter
  10. Wikipedia. "Lighthill report." https://en.wikipedia.org/wiki/Lighthill_report
  11. Wikipedia. "Expert system." https://en.wikipedia.org/wiki/Expert_system
  12. Wikipedia. "2024 Nobel Prizes." https://en.wikipedia.org/wiki/2024_Nobel_Prizes
  13. Wikipedia. "TD-Gammon." https://en.wikipedia.org/wiki/TD-Gammon
  14. Wikipedia. "Deep Blue versus Garry Kasparov." https://en.wikipedia.org/wiki/Deep_Blue_versus_Garry_Kasparov
  15. Wikipedia. "IBM Watson." https://en.wikipedia.org/wiki/IBM_Watson
  16. Wikipedia. "ImageNet." https://en.wikipedia.org/wiki/ImageNet
  17. Wikipedia. "AlexNet." https://en.wikipedia.org/wiki/AlexNet
  18. Wikipedia. "Turing Award." https://en.wikipedia.org/wiki/Turing_Award
  19. Wikipedia. "AlphaGo." https://en.wikipedia.org/wiki/AlphaGo
  20. Wikipedia. "AlphaGo versus Lee Sedol." https://en.wikipedia.org/wiki/AlphaGo_versus_Lee_Sedol
  21. Vaswani, A. et al. "Attention Is All You Need." arXiv, June 12, 2017. https://arxiv.org/abs/1706.03762
  22. Wikipedia. "Generative pre-trained transformer." https://en.wikipedia.org/wiki/Generative_pre-trained_transformer
  23. Wikipedia. "BERT (language model)." https://en.wikipedia.org/wiki/BERT_(language_model)
  24. Wikipedia. "GPT-3." https://en.wikipedia.org/wiki/GPT-3
  25. Wikipedia. "AlphaFold." https://en.wikipedia.org/wiki/AlphaFold
  26. Wikipedia. "GPT-4." https://en.wikipedia.org/wiki/GPT-4
  27. Wikipedia. "Claude (language model)." https://en.wikipedia.org/wiki/Claude_(language_model)
  28. Wikipedia. "Llama (language model)." https://en.wikipedia.org/wiki/Llama_(language_model)
  29. Wikipedia. "Gemini (language model)." https://en.wikipedia.org/wiki/Gemini_(language_model)
  30. Wikipedia. "AI boom." https://en.wikipedia.org/wiki/AI_boom
  31. Wikipedia. "AI Safety Summit." https://en.wikipedia.org/wiki/AI_Safety_Summit
  32. Wikipedia. "Artificial Intelligence Act." https://en.wikipedia.org/wiki/Artificial_Intelligence_Act
  33. Wikipedia. "OpenAI o1." https://en.wikipedia.org/wiki/OpenAI_o1
  34. Wikipedia. "DeepSeek." https://en.wikipedia.org/wiki/DeepSeek
  35. Wikipedia. "Stargate LLC." https://en.wikipedia.org/wiki/Stargate_LLC
  36. Wikipedia. "GPT-5." https://en.wikipedia.org/wiki/GPT-5
  37. Wikipedia. "India AI Impact Summit." https://en.wikipedia.org/wiki/India_AI_Impact_Summit

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

Suggest edit