How did we get to ChatGPT? The inventions, ideas and rules that made today’s AI, from the first electronic switches to reasoning models. Follow what each milestone built on and what it led to, right up to our news.
100 milestones · 1904 to today · each checked in at least 2 sources
When it happened
Each square is one decade. The brighter the square, the more milestones of that kind: progress was slow for decades, then sped up. Choose a theme to see only its milestones, or a decade to jump to it.
Milestones per theme per decade, from the 1900s to the 2020s
Neurophysiologist Warren McCulloch and logician Walter Pitts described nerve cells as simple on-off units with a firing threshold. They showed that networks of such units can carry out logical calculations.
Why it matters Their neuron model became the starting point for artificial neural networks and influenced John von Neumann’s thinking about computers.
Note Widely credited as the first mathematical model of an artificial neuron. The Computer History Museum notes that the model neurons were much simpler than real nerve cells.
In 1948, MIT mathematician Norbert Wiener published “Cybernetics: Or Control and Communication in the Animal and the Machine”. It named the study of control and communication after a Greek word for steersman.
Why it matters It helped spread the idea of feedback to engineering, biology and computing, and it lies behind words like cyberspace.
Note Others had advanced related ideas, partly independently, by 1948. Cybernetics later faded as a separate field as its ideas moved into other fields.
In 1949, psychologist Donald Hebb published “The Organization of Behavior”. It proposed that when one brain cell repeatedly helps to fire another, the link between the two cells grows stronger.
Why it matters The idea of learning by strengthening links still influences neuroscience, robotics and computer science.
Note Hebb said the rule was only a clearer statement of a widely held idea. The slogan “cells that fire together, wire together” is a later summary, not Hebb’s own words.
In October 1950, Alan Turing published “Computing Machinery and Intelligence” in the journal Mind. It replaced the question “Can machines think?” with a test called the imitation game.
Why it matters The imitation game, later called the Turing test, became a lasting way to frame the question of machine intelligence.
Note The game starts with a man, a woman and an interrogator; the machine takes one player’s place. Scholars still debate what passing the test would prove.
In 1951, graduate students Marvin Minsky and Dean Edmonds built SNARC, a machine of 40 electronic units that imitated linked brain cells. When rewarded, it strengthened the connections that had just been active, like a rat learning a maze.
Why it matters Later writers call it the first neural network built in hardware, and it explored learning by reward long before today’s machine learning.
Note Minsky’s own 2011 recollection hesitates between 1951 and 1952 before settling on summer 1951. Sources disagree on where it was built and who paid for it.
In 1952, Arthur Samuel of IBM got a checkers program running on the IBM 701 computer. Later versions improved by learning from recorded games and from games against themselves.
Why it matters It was an early example of a computer improving through experience, and it made checkers a classic test for machine learning.
Note Samuel’s own note dates the first working 701 program to 1952. A version that could learn was finished in 1955 and shown on television in February 1956.
On 31 August 1955, John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon proposed a summer study of “artificial intelligence” at Dartmouth College. McCarthy is credited with choosing the name.
Why it matters The name stuck, and the proposal’s bold idea that machines can simulate intelligence became a shared vision for the field.
Note McCarthy said in 2007 that the phrase may have been used before, but the source was never found. Some histories date the name to the 1956 meeting; the written proposal is from August 1955.
In summer 1956, researchers came in turns to Dartmouth College in New Hampshire to work on making machines simulate intelligence. John McCarthy led the project with Marvin Minsky, Nathaniel Rochester and Claude Shannon.
Why it matters It brought early AI researchers together and set the goal of machines that simulate intelligence, even though they did not agree on one method.
Note Often taken as the start of AI as a research field. Attendees came at different times, worked mostly on their own projects and did not agree on a common theory; the sources agree only on summer 1956.
Allen Newell, Herbert A. Simon and Cliff Shaw built Logic Theorist, a program that finds proofs in symbolic logic. Shown at the Dartmouth workshop in 1956, it proved 38 theorems from Whitehead and Russell’s Principia Mathematica.
Why it matters It showed that a computer could use rules of thumb, called heuristics, to search for solutions, an idea that shaped early AI.
Note Often called the first AI program; Carnegie Mellon’s own pages say both “the first” and “one of the first”. Work began in 1955, and the first papers appeared in 1956.
At the Cornell Aeronautical Laboratory, Frank Rosenblatt designed the perceptron, a one-layer neural network that adjusts itself when it makes mistakes. In July 1958, the US Navy showed an IBM 704 teaching itself to tell cards marked on the left from cards marked on the right.
Why it matters It was an early machine that learned from its own mistakes, and it helped lay the ground for today’s deep learning.
Note Often called the first neural network, but earlier neuron models and Minsky’s SNARC came first. The idea first appeared in a January 1957 report; the Mark I perceptron hardware was built after the 1958 demonstration.
In July 1959, Arthur Samuel of IBM published a paper on teaching a computer to play checkers by learning from experience. Running on an IBM 704, the program learned to play better than its own programmer after 8 to 10 hours of play.
Why it matters It showed that a computer can improve by experience, and “machine learning” later became the name of a whole field.
Note Many sources say Samuel coined “machine learning” here, but the paper uses the phrase as if it already existed. The famous line about learning “without being explicitly programmed” is not in the paper.
In 1965, Edward Feigenbaum and Joshua Lederberg began the DENDRAL project at Stanford University. The program helped chemists find the structure of unknown molecules from mass spectrometer data.
Why it matters DENDRAL showed that stored expert knowledge, not clever tricks, makes a program effective, and it inspired later expert systems.
Note Widely called the first expert system. Lederberg’s first DENDRAL reports are from 1964, and the Stanford project with Feigenbaum began in spring 1965.
In 1969, Marvin Minsky and Seymour Papert published the book “Perceptrons”. It used mathematics to show what simple one-layer neural networks cannot do.
Why it matters The book discouraged some researchers from studying neural networks, and interest in them grew again only in later decades.
Note How much the book caused the fall of neural network research is debated. Cornell says it sealed the fate of Rosenblatt’s work; a survey says it discouraged some researchers.
In 1980, Digital Equipment Corporation began using R1, later called XCON, to configure orders for its VAX computers. John McDermott of Carnegie Mellon University built it from about 770 if-then rules.
Why it matters It showed that a rule-based program could do a real job in a company, not just in a lab.
Note Work began in December 1978, DEC began using R1 in 1980, and the full journal paper appeared in 1982, so other lists may give a different year.
In 1980, Kunihiko Fukushima published the neocognitron, a layered neural network for recognising visual patterns. It learned without a teacher and recognised shapes even when they moved or were slightly distorted.
Why it matters Its stacked layers of feature detectors are seen as an early form of the convolutional networks used in image recognition today.
Note Fukushima first reported the idea in 1979, in Japanese and at a conference, and the prize pages give 1979. The full English paper is from April 1980.
In April 1982, physicist John Hopfield described a network of simple on-off neurons that works as a memory. Given part of a stored pattern, or a damaged copy, the network settles on the whole pattern.
Why it matters Geoffrey Hinton and colleagues built on it for the Boltzmann machine, and Hopfield shared the 2024 Nobel Prize in Physics for this work.
Note Similar network models already existed, for example by Shun’ichi Amari in 1972 and William Little in 1974. The Nobel background names Hopfield’s paper as the key memory model.
At the AAAI conference in Austin, Texas, in 1984, a panel of AI researchers warned that hopes for AI ran too high. Drew McDermott described a possible collapse of AI funding that some were calling an “AI winter”.
Why it matters It came true: interest and funding fell from the mid-1980s, and cheaper workstations undermined the companies making special AI computers, called Lisp machines.
Note This period is often called the second AI winter, but sources disagree on when it began: AI100 says the mid-1980s, others the late 1980s. The panel is the earliest use of the term we could confirm.
In Nature, David Rumelhart, Geoffrey Hinton and Ronald Williams described backpropagation, a way to train neural networks with several layers by sending errors backwards. They showed that hidden layers can learn useful features of a task.
Why it matters It showed how to train networks with hidden layers, and backpropagation later trained deep networks such as LeCun’s.
Note The method is older: a survey traces it to Seppo Linnainmaa (1970) and its first use in neural networks to Paul Werbos (1981). The Nobel committee says the authors reinvented it; the 1986 paper made it popular.
Yann LeCun and colleagues at Bell Labs trained a neural network with backpropagation to read handwritten zip codes from US mail. Its design built in knowledge about images, so one network handled the whole job.
Why it matters Its design, later called a convolutional network, was used by several US banks to read handwritten cheques from the mid-1990s.
Note Fukushima’s neocognitron (1980) was a similar design. The new step was training such a network with backpropagation on real mail.
At IBM, Gerald Tesauro built a neural network that taught itself backgammon by playing against itself. In 1992, its second version played 38 games against top players and lost by only seven points in total.
Why it matters It showed that a program could learn to play near the level of the best humans from its own games.
Note Sources give different numbers of training games for the early version, so we give none.
Corinna Cortes and Vladimir Vapnik, at AT&T Bell Labs, described the support-vector network for sorting data into two groups. It worked even when the groups overlapped, and was tested on reading handwritten digits.
Why it matters It became a widely used method for analysing text, images and other data.
Note Support-vector ideas are older: Burges traces them to Vapnik’s work in the late 1970s, and the kernel idea was shown in 1992. The 1995 paper handled data that cannot be split cleanly.
On 11 May 1997, IBM’s Deep Blue won its six-game match against world chess champion Garry Kasparov in New York, 3.5 to 2.5. It is widely credited as the first computer to beat a reigning world champion in a match.
Why it matters It reached a goal pursued since the late 1940s and was widely seen as a symbolic contest between machine and human intelligence.
Note Kasparov accused the IBM team of cheating, IEEE Spectrum reports. Deep Blue had won one game against Kasparov in February 1996, but Kasparov won that match 4–2.
Sepp Hochreiter and Jürgen Schmidhuber described long short-term memory (LSTM), a recurrent neural network built from memory cells controlled by gates. In tests, it linked events more than 1,000 steps apart, which earlier recurrent networks struggled with.
Why it matters It let neural networks learn from long sequences and became an important method for speech and language.
Note The method was first described in a 1995 technical report, and Hochreiter’s 1991 thesis analysed the underlying problem. The date shown is the journal paper, November 1997.
In July 2006, Geoffrey Hinton, Simon Osindero and Yee-Whye Teh published a fast way to train neural networks with many layers. It trains one layer at a time, then fine-tunes the whole network.
Why it matters It made networks with many layers trainable and was a milestone on the road to what is now called deep learning.
Note Deep networks were studied before 2006: a survey by Jürgen Schmidhuber traces them back to 1965 and says the name “deep learning” took off around 2006.
Demis Hassabis, Shane Legg and Mustafa Suleyman started DeepMind in London in 2010. Its stated aim was to “solve intelligence”, using ideas from machine learning and neuroscience.
Why it matters Google bought DeepMind in 2014, and the lab went on to build AlphaGo and AlphaFold.
Note DeepMind says it started in 2010, and the UK company register shows it was registered in September 2010 and renamed DeepMind Technologies in November 2010. MIT Technology Review (2014) says 2011.
A deep neural network built by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton won the 2012 ImageNet image-recognition contest. It missed the right answer in its top five guesses 15.3% of the time, against 26.2% for the best rival.
Why it matters AlexNet showed that deep networks, GPUs and big data sets could beat older methods, and neural networks soon took over computer vision.
Note Contest entries closed on 30 September 2012 and the results came out in October 2012. The paper followed in December 2012 at the NIPS conference.
Ian Goodfellow and colleagues at the Université de Montréal described a new way to train models that create data, such as images. Two networks compete: a generator makes fakes, and a discriminator tries to spot them.
Why it matters GANs set off a wave of research on realistic generated images, and also raised fears about convincing fakes.
In December 2015, Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun at Microsoft Research described residual networks with up to 152 layers. A group of them won the 2015 ImageNet classification contest, with 3.57% error.
Why it matters Shortcut connections made it possible to train networks with far more layers, and the transformer later used them too.
Note First posted on 10 December 2015, the day the ImageNet 2015 results came out. It won the best paper award at the CVPR conference in June 2016.
OpenAI was announced on 11 December 2015 as a non-profit AI research company. Backers including Elon Musk and Sam Altman pledged $1 billion, and Ilya Sutskever became research director.
Why it matters It created a well-funded lab that promised to share its research openly; it later built the GPT models and ChatGPT.
Note Reports describe the $1 billion differently: TechCrunch says “contributed”, The Christian Science Monitor says “investment”.
DeepMind’s AlphaFold 2 took part in CASP14, a contest to predict the 3D shapes of proteins, and scored far ahead of the other teams. The organisers said about two thirds of its predictions matched lab-quality results.
Why it matters DeepMind published the method in Nature in July 2021, and scientists said it could speed up biology and drug research.
Note DeepMind called it a solution to a 50-year-old problem. CASP’s organisers noted limits: single proteins, not complexes. We do not say it “solved” protein folding.
OpenAI showed DALL·E, a neural network that creates images from written captions, such as an armchair shaped like an avocado. Like GPT-3, it is a transformer, trained on 250 million pairs of images and text.
Why it matters It showed that a GPT-style model can turn words into pictures; DALL·E 2 followed in April 2022, using a different method.
On 22 August 2022, Stability AI released Stable Diffusion, a model that turns text prompts into images. The code and model files were free to download, under a licence with rules against misuse.
Why it matters Anyone with a good graphics card could run and adapt a capable image model, which also sparked debate about artists’ work in training data.
Note This was the public release; researchers had received the model earlier. “Open” here means downloadable model files under a licence with use rules, not classic open source.
On 12 September 2024, OpenAI released o1-preview and o1-mini, models that work through a long chain of thought before they answer. OpenAI said they learned this by large-scale reinforcement learning.
Why it matters OpenAI reported large gains on maths and coding tests, and four months later DeepSeek-R1 claimed to match it.
Note Whether these models really reason is debated. Asking models to write out their steps was shown earlier, in a January 2022 paper, so we don’t call o1 the first.
The 2024 Nobel Prize in Physics went to John Hopfield and Geoffrey Hinton for work on artificial neural networks. The Chemistry prize went to David Baker for protein design, and to Demis Hassabis and John Jumper for protein structure prediction.
Why it matters Two science Nobels in two days honoured AI work; the Academy said AlphaFold2 had predicted about 200 million protein structures.
Note The two prizes were announced on different days, Physics on 8 October and Chemistry on 9 October 2024, so only the month is given.
On 20 January 2025, the Chinese company DeepSeek released DeepSeek-R1 with open weights under an MIT licence. DeepSeek said it matched OpenAI’s o1 on maths, code and reasoning tasks.
Why it matters It put a strong reasoning model in everyone’s hands for free, and rivals such as Alibaba and OpenAI reacted quickly.
Note Not the first open reasoning model: IEEE Spectrum names Alibaba’s QwQ as an earlier one. A web-only preview, R1-Lite, came on 20 November 2024.
A milestone goes in when it changed what AI can do, the tools AI is built with, how many people use AI, or the rules for AI. It must be at least a year old: newer events are news.
We checked every milestone in at least two sources and read them: the original paper, patent, announcement or law where it exists, and a history source such as a museum, an encyclopedia or a report from the time. We never cite Wikipedia. Many “firsts” are disputed, so we say when they are, and dates are only as exact as the sources agree.
Last checked 4 October 2026. Found a mistake? Tell us, and we will correct it. How we report