How did we get to ChatGPT? The inventions, ideas and rules that made today’s AI, from the first electronic switches to reasoning models. Follow what each milestone built on and what it led to, right up to our news.
100 milestones · 1904 to today · each checked in at least 2 sources
When it happened
Each square is one decade. The brighter the square, the more milestones of that kind: progress was slow for decades, then sped up. Choose a theme to see only its milestones, or a decade to jump to it.
Milestones per theme per decade, from the 1900s to the 2020s
John Ambrose Fleming built a glass bulb with a hot filament and a metal plate that lets current pass one way only. The bulb could detect radio signals, and Fleming filed a patent for this “oscillation valve” on 16 November 1904.
Why it matters It detected Morse-code radio signals right away and laid the foundation for electronics.
Note Fleming had studied current flow in light bulbs (the Edison effect) since the 1880s, so the idea was not new in 1904. The patent was filed on 16 November 1904; a Science Museum record says Fleming used a valve in October 1904.
American inventor Lee De Forest put a wire grid between the filament and the plate of a vacuum tube and named it the Audion. A small signal on the grid controlled a larger current, so the tube could amplify weak signals.
Why it matters It became the first widely available electronic amplifier and changed telephone and radio technology.
Note De Forest showed his first Audion in October 1906; the grid patent was filed in January 1907 and describes a radio detector. Sources disagree on whether De Forest knew Fleming’s valve. Robert von Lieben patented a different amplifying tube in 1906.
Karel Čapek’s play R.U.R. (Rossum’s Universal Robots), about factory-made workers who rebel, opened in January 1921. Karel Čapek later credited the word “robot” to Josef Čapek, Karel’s brother; it comes from the Czech robota, meaning forced labour.
Why it matters Staged from New York to Tokyo within a few years, the play spread the word “robot” into many languages.
Note The first performance was on 2 January 1921, by an amateur group in Hradec Králové; the Prague National Theatre premiere followed on 25 January 1921. The play was printed in November 1920.
Kurt Gödel proved that any rule system without contradictions, if strong enough for basic arithmetic, has statements it can neither prove nor disprove. Such a system also cannot prove that it is free of contradictions.
Why it matters It badly damaged David Hilbert’s plan to prove mathematics complete and free of contradictions, and it shaped Alan Turing’s work on computation.
Note Sources give January 1931 (Stanford Encyclopedia) or December 1931 (the journal issue). Gödel first announced the first theorem on 7 September 1930, at a conference in Königsberg.
Alan Turing described an imaginary machine that reads and writes symbols on a tape by simple rules. Turing also proved that one “universal” machine can imitate any other such machine.
Why it matters It defined computation precisely, showed that some questions no machine can answer, and its universal machine is the idea behind programmable computers.
Note The paper was received on 28 May 1936 and read on 12 November 1936; the journal volume is dated 1936–37. Alonzo Church reached a matching answer slightly earlier, by a different method.
In an MIT master’s thesis, Claude Shannon showed that George Boole’s algebra of logic can describe networks of relays and switches. The thesis also shows how to design circuits for tasks such as adding binary numbers.
Why it matters Engineers could design and check switching circuits with algebra instead of trial and error, which became a base for digital computer design.
Note MIT’s catalogue dates the thesis to 1940, the year of the degree; the thesis itself is signed 10 August 1937. Victor Shestakov in the USSR proposed similar ideas in 1935 but published them only in 1941.
At Iowa State College, John Vincent Atanasoff and graduate student Clifford Berry built a prototype in 1939 and a full-size machine by 1942. It solved systems of linear equations, using binary numbers and vacuum tubes.
Why it matters It pioneered electronic, binary computing, and in 1973 a US court used it to void the ENIAC patent.
Note Often called the first electronic digital computer, but it solved one type of problem and could not store programs. The IEEE plaque says the prototype was built in October 1939; the Computer History Museum says work began in 1938.
German engineer Konrad Zuse showed the Z3 to a group of scientists in Berlin on 12 May 1941. The machine used electric relays, worked in binary and ran programs punched on film tape.
Why it matters It showed that a programmable, binary machine could work, and the Deutsches Museum dates the start of the universal-computer age to it.
Note German museums call it the first computer; the Computer History Museum calls it an early one. It used relays, not electronics. 12 May 1941 is the day it was shown to scientists.
Neurophysiologist Warren McCulloch and logician Walter Pitts described nerve cells as simple on-off units with a firing threshold. They showed that networks of such units can carry out logical calculations.
Why it matters Their neuron model became the starting point for artificial neural networks and influenced John von Neumann’s thinking about computers.
Note Widely credited as the first mathematical model of an artificial neuron. The Computer History Museum notes that the model neurons were much simpler than real nerve cells.
Engineer Tommy Flowers and a Post Office team built Colossus, an electronic machine with vacuum tubes, to break German Lorenz messages. The first one was working at Bletchley Park by early February 1944; ten were built in all.
Why it matters It cut the time to break Lorenz messages from weeks to hours, but secrecy kept it out of computing history for decades.
Note Sources call it the first large-scale electronic digital computer or the first programmable one, but its programming was limited and it did not store programs. It arrived between late December 1943 and January 1944.
In 1945, John von Neumann wrote a report on the planned EDVAC computer. It described a machine that would keep its instructions in memory.
Why it matters Copies spread widely, and the design guided later stored-program computers, often called von Neumann machines.
Note Credit is disputed: only von Neumann’s name is on the report, but Eckert, Mauchly and other Moore School staff said the ideas were shared. The title page says 30 June 1945; one history says May 1945.
In July 1945, Vannevar Bush published “As We May Think” in The Atlantic Monthly. It described the memex, a desk that held books and records on microfilm and let a user link items into trails.
Why it matters Its idea of linked trails of information influenced later computer pioneers such as Douglas Engelbart.
Note The memex was a proposal for a future device, not a working machine. Engelbart read a reprint in 1945, but the essay shaped Engelbart’s own research only from about 1959.
ENIAC was unveiled on 14 February 1946. Built at the University of Pennsylvania for the US Army, it used about 18,000 vacuum tubes to calculate artillery tables.
Why it matters It showed that large electronic computers could work and influenced the stored-program machines that followed.
Note Often called the first general-purpose electronic computer, but the Stanford Encyclopedia of Philosophy says the earlier Colossus was similar and ENIAC was far from general-purpose. The New York Times front-page story ran on 15 February 1946.
In December 1947, John Bardeen and Walter Brattain built a point-contact transistor from germanium at Bell Labs. Their group, led by William Shockley, made a device that amplified signals, widely credited as the first transistor.
Why it matters Transistors replaced vacuum tubes and relays, and later made integrated circuits and the information age possible.
Note Sources name different days: 16 December (the first working amplifier) and 23 December (the demonstration to Bell Labs managers). The public announcement came on 30 June 1948. Julius Lilienfeld had patented field-effect designs from 1925.
In 1948, MIT mathematician Norbert Wiener published “Cybernetics: Or Control and Communication in the Animal and the Machine”. It named the study of control and communication after a Greek word for steersman.
Why it matters It helped spread the idea of feedback to engineering, biology and computing, and it lies behind words like cyberspace.
Note Others had advanced related ideas, partly independently, by 1948. Cybernetics later faded as a separate field as its ideas moved into other fields.
On 21 June 1948, the Manchester Baby ran its first program. It is widely credited as the first computer to run a program stored in its own electronic memory.
Why it matters It proved that the stored-program idea and the new Williams–Kilburn memory tube worked, and it led to the Ferranti Mark 1.
Note IEEE words the claim carefully: the first to run a program from addressable read-write electronic memory. The Stanford Encyclopedia of Philosophy says the roles of Max Newman and Alan Turing have been neglected.
In July 1948, Claude Shannon published “A Mathematical Theory of Communication” in the Bell System Technical Journal. It showed how to measure information in bits and how much a noisy channel can carry.
Why it matters It became the foundation of digital communication and today shapes systems from compact discs to deep-space probes.
Note The paper was printed in two parts, in July and October 1948; the date given is the first part. John W. Tukey suggested the word “bit”.
In 1949, psychologist Donald Hebb published “The Organization of Behavior”. It proposed that when one brain cell repeatedly helps to fire another, the link between the two cells grows stronger.
Why it matters The idea of learning by strengthening links still influences neuroscience, robotics and computer science.
Note Hebb said the rule was only a clearer statement of a widely held idea. The slogan “cells that fire together, wire together” is a later summary, not Hebb’s own words.
In October 1950, Alan Turing published “Computing Machinery and Intelligence” in the journal Mind. It replaced the question “Can machines think?” with a test called the imitation game.
Why it matters The imitation game, later called the Turing test, became a lasting way to frame the question of machine intelligence.
Note The game starts with a man, a woman and an interrogator; the machine takes one player’s place. Scholars still debate what passing the test would prove.
In 1951, two makers began delivering computers to customers. Ferranti’s first Mark 1 reached Manchester University in February, and the first UNIVAC I went to the US Census Bureau.
Why it matters Buyers could now purchase a ready-made computer instead of building one, and UNIVAC I brought computers into public view.
Note Ferranti’s first Mark 1 arrived in February 1951; the Census Bureau gives 31 March 1951 for the UNIVAC I contract and 14 June 1951 for its dedication. “First” depends on the definition: the Computer History Museum says Ferranti “probably” holds the title.
In 1951, graduate students Marvin Minsky and Dean Edmonds built SNARC, a machine of 40 electronic units that imitated linked brain cells. When rewarded, it strengthened the connections that had just been active, like a rat learning a maze.
Why it matters Later writers call it the first neural network built in hardware, and it explored learning by reward long before today’s machine learning.
Note Minsky’s own 2011 recollection hesitates between 1951 and 1952 before settling on summer 1951. Sources disagree on where it was built and who paid for it.
In 1952, Arthur Samuel of IBM got a checkers program running on the IBM 701 computer. Later versions improved by learning from recorded games and from games against themselves.
Why it matters It was an early example of a computer improving through experience, and it made checkers a classic test for machine learning.
Note Samuel’s own note dates the first working 701 program to 1952. A version that could learn was finished in 1955 and shown on television in February 1956.
On 7 January 1954, Georgetown University and IBM showed a computer turning more than sixty Russian sentences into English in New York. The IBM 701 used a vocabulary of only 250 words and six grammar rules.
Why it matters It showed machine translation working on a real computer and won funding, but it made good translation seem closer than it was.
Note Widely described as the first public demonstration of machine translation. It was a showcase: the words and rules were chosen for a small set of sentences, so it could not translate general text.
On 31 August 1955, John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon proposed a summer study of “artificial intelligence” at Dartmouth College. McCarthy is credited with choosing the name.
Why it matters The name stuck, and the proposal’s bold idea that machines can simulate intelligence became a shared vision for the field.
Note McCarthy said in 2007 that the phrase may have been used before, but the source was never found. Some histories date the name to the 1956 meeting; the written proposal is from August 1955.
In summer 1956, researchers came in turns to Dartmouth College in New Hampshire to work on making machines simulate intelligence. John McCarthy led the project with Marvin Minsky, Nathaniel Rochester and Claude Shannon.
Why it matters It brought early AI researchers together and set the goal of machines that simulate intelligence, even though they did not agree on one method.
Note Often taken as the start of AI as a research field. Attendees came at different times, worked mostly on their own projects and did not agree on a common theory; the sources agree only on summer 1956.
Allen Newell, Herbert A. Simon and Cliff Shaw built Logic Theorist, a program that finds proofs in symbolic logic. Shown at the Dartmouth workshop in 1956, it proved 38 theorems from Whitehead and Russell’s Principia Mathematica.
Why it matters It showed that a computer could use rules of thumb, called heuristics, to search for solutions, an idea that shaped early AI.
Note Often called the first AI program; Carnegie Mellon’s own pages say both “the first” and “one of the first”. Work began in 1955, and the first papers appeared in 1956.
A small IBM team led by John Backus built FORTRAN, a language that let scientists write mathematical formulas instead of machine code. A compiler translated the formulas, and the first version reached IBM 704 users in April 1957.
Why it matters It showed that compiled code could run nearly as fast as hand-written code, which made programming faster, cheaper and open to more people.
Note Not the first compiler: Backus credited an MIT system by Laning and Zierler (1954) as the first working algebraic compiler. The manual is dated October 1956; the compiler reached users in April 1957.
At the Cornell Aeronautical Laboratory, Frank Rosenblatt designed the perceptron, a one-layer neural network that adjusts itself when it makes mistakes. In July 1958, the US Navy showed an IBM 704 teaching itself to tell cards marked on the left from cards marked on the right.
Why it matters It was an early machine that learned from its own mistakes, and it helped lay the ground for today’s deep learning.
Note Often called the first neural network, but earlier neuron models and Minsky’s SNARC came first. The idea first appeared in a January 1957 report; the Mark I perceptron hardware was built after the 1958 demonstration.
In autumn 1958, John McCarthy and Marvin Minsky started an AI project at MIT, where work began on LISP, a programming language for symbols. A 1960 paper described the system, which ran on the IBM 704.
Why it matters LISP introduced ideas such as automatic memory clean-up and recursive functions, and it became a long-lasting language for AI programs.
Note McCarthy says the main ideas came together between 1956 and 1958, and building LISP began in autumn 1958. The first full paper appeared in April 1960.
On 12 September 1958, Jack Kilby of Texas Instruments showed a working circuit built on one bar of germanium. In July 1959, Robert Noyce of Fairchild filed a patent for a silicon chip design that could be made in large numbers.
Why it matters It ended the wiring of thousands of parts by hand and made later electronics much smaller, cheaper and more powerful.
Note IEEE and others credit Kilby with the first working integrated circuit, but several labs were working on the idea. Noyce’s silicon design was the practical form, and both are called co-inventors.
In July 1959, Arthur Samuel of IBM published a paper on teaching a computer to play checkers by learning from experience. Running on an IBM 704, the program learned to play better than its own programmer after 8 to 10 hours of play.
Why it matters It showed that a computer can improve by experience, and “machine learning” later became the name of a whole field.
Note Many sources say Samuel coined “machine learning” here, but the paper uses the phrase as if it already existed. The famous line about learning “without being explicitly programmed” is not in the paper.
In 1961, the first Unimate robot arm went to work at a General Motors plant in New Jersey. It lifted hot die-cast metal parts from a machine and stacked them.
Why it matters Unimate is widely called the first industrial robot, and it opened the way for programmable robot arms in factories.
Note “First industrial robot” is a common label, not a settled record. The sources give only the year 1961, not the month.
In 1965, Edward Feigenbaum and Joshua Lederberg began the DENDRAL project at Stanford University. The program helped chemists find the structure of unknown molecules from mass spectrometer data.
Why it matters DENDRAL showed that stored expert knowledge, not clever tricks, makes a program effective, and it inspired later expert systems.
Note Widely called the first expert system. Lederberg’s first DENDRAL reports are from 1964, and the Stanford project with Feigenbaum began in spring 1965.
On 19 April 1965, Gordon Moore wrote in Electronics magazine that the number of parts on a chip would double about every year. Moore said falling costs could put as many as 65,000 parts on one chip by 1975.
Why it matters The idea, later called Moore’s law, set the pace of the chip industry; in 1975 Moore slowed the forecast to every two years.
From 1966 to 1972, the AI Center at SRI built Shakey, a mobile robot with a TV camera and other sensors. Shakey could plan tasks, find routes and move simple objects.
Why it matters Shakey’s planning and route-finding methods, such as STRIPS and A* search, shaped later robots and AI software.
Note SRI and IEEE call Shakey the first mobile robot that could reason about its surroundings. Simpler mobile robots came earlier, such as Grey Walter’s robot tortoises around 1950.
In January 1966, Joseph Weizenbaum of MIT published a paper on ELIZA, a program that chats in everyday English. Its best-known script acted like a psychotherapist, reacting to keywords in what the user typed.
Why it matters ELIZA showed how easily people believe a program understands them, and Weizenbaum later warned against handing human decisions to machines.
Note MIT News calls ELIZA “perhaps the first” chatbot. January 1966 is the date of the paper; the program existed before then.
In 1969, Marvin Minsky and Seymour Papert published the book “Perceptrons”. It used mathematics to show what simple one-layer neural networks cannot do.
Why it matters The book discouraged some researchers from studying neural networks, and interest in them grew again only in later decades.
Note How much the book caused the fall of neural network research is debated. Cornell says it sealed the fate of Rosenblatt’s work; a survey says it discouraged some researchers.
On 29 October 1969, a computer at UCLA sent the first message over the ARPANET to a computer at SRI. It was meant to be “LOGIN”, but the SRI computer crashed after “LO”; the full login worked within an hour.
Why it matters It joined the first two nodes of the ARPANET, the research network that later grew into the internet.
Note Often called the first message on the internet. The ARPANET was an early network that later became part of the internet.
In November 1971, Intel announced the 4004, widely credited as the first commercial microprocessor: a whole processor on one chip. It began as a chip set for a Busicom calculator.
Why it matters It made computing cheaper and smaller, a step Intel says helped start the modern information age.
Note “First” depends on the definition: the MP944 chip set (about 1970, for the F-14 jet) and Four-Phase’s AL1 (1969) came earlier. IEEE Spectrum credits the 4004 if “microprocessor” means a programmable CPU on one chip.
In 1972, Alain Colmerauer and Philippe Roussel built the first Prolog system at the University of Aix-Marseille in France. It grew from a question-answering project and Robert Kowalski’s ideas about programming in logic.
Why it matters Prolog spread to many universities and became a leading language for logic programming and early AI research.
Note Sources agree on 1972 but differ on the season (summer or autumn). Some give 1971, because of earlier work on logic and parsing.
In 1973, Waseda University in Tokyo completed WABOT-1, a human-shaped robot. It could walk on two legs, grip objects and talk with people in simple Japanese.
Why it matters Waseda says this early success helped start modern humanoid robot research and shaped many later robots.
Note Waseda and others call WABOT-1 the first full-scale humanoid robot; this is a label, not a measured record. The sources give the year but no exact day.
Mathematician James Lighthill reviewed AI research for Britain’s Science Research Council. The report, published in 1973, said AI had not delivered what was promised, and it was harshest on robot building.
Why it matters The council then reorganised its AI funding, and the report is often blamed for the first “AI winter”.
Note The report is dated July 1972 but was published in early 1973. Historian Jon Agar shows that the council steered the review, so blaming Lighthill alone for the first AI winter is too simple.
In 1977, three ready-made computers went on sale: the Apple II, the Commodore PET and the TRS-80. Unlike earlier kits, they came fully assembled for ordinary buyers.
Why it matters They brought computers into homes and schools, and were among the first personal computers sold to a mass market.
Note The three reached buyers at different times in 1977: the Apple II shipped on 10 June and the TRS-80 went on sale on 3 August. Sources disagree on which came first.
In 1979, the Stanford Cart, a small robot with a TV camera, crossed a room with obstacles, such as a chair, without human help. It moved about one metre at a time, then paused for 10 to 15 minutes to look and plan.
Why it matters It showed that a robot could find its way around obstacles from camera pictures alone, but only very slowly.
Note The cart itself dates from the early 1960s; Moravec rebuilt it with stereo vision in 1977. Popular accounts say “a chair-filled room”, but the thesis describes courses with a chair and cardboard shapes.
In 1980, Digital Equipment Corporation began using R1, later called XCON, to configure orders for its VAX computers. John McDermott of Carnegie Mellon University built it from about 770 if-then rules.
Why it matters It showed that a rule-based program could do a real job in a company, not just in a lab.
Note Work began in December 1978, DEC began using R1 in 1980, and the full journal paper appeared in 1982, so other lists may give a different year.
In 1980, Kunihiko Fukushima published the neocognitron, a layered neural network for recognising visual patterns. It learned without a teacher and recognised shapes even when they moved or were slightly distorted.
Why it matters Its stacked layers of feature detectors are seen as an early form of the convolutional networks used in image recognition today.
Note Fukushima first reported the idea in 1979, in Japanese and at a conference, and the prize pages give 1979. The full English paper is from April 1980.
On 12 August 1981, IBM presented its Personal Computer in New York. Don Estridge’s team in Boca Raton built it around an Intel 8088 chip and an operating system from Microsoft.
Why it matters IBM used outside parts and published its design, so rivals built compatible “clones” and the PC became widely used in business.
In April 1982, physicist John Hopfield described a network of simple on-off neurons that works as a memory. Given part of a stored pattern, or a damaged copy, the network settles on the whole pattern.
Why it matters Geoffrey Hinton and colleagues built on it for the Boltzmann machine, and Hopfield shared the 2024 Nobel Prize in Physics for this work.
Note Similar network models already existed, for example by Shun’ichi Amari in 1972 and William Little in 1974. The Nobel background names Hopfield’s paper as the key memory model.
In October 1981, Japan announced a ten-year national plan for “fifth generation” computers built on logic programming. The project formally began in April 1982, when the ICOT institute opened in Tokyo.
Why it matters The plan alarmed the US and other countries, which funded their own AI projects; it was later judged a failure.
Note Japan announced the plan in October 1981, but ICOT and the project started in April 1982, so sources give 1981 or 1982. Calling it a failure is a later judgement.
On 1 January 1983, every computer on the ARPANET switched from the older protocol, NCP, to TCP/IP. A 1981 plan by Jon Postel set the deadline, and the change went surprisingly smoothly.
Why it matters TCP/IP let many separate networks link up under shared rules: the basis of the internet that followed.
Note The date is firm, but the move was staged: under RFC 801, computers ran both protocols during 1982.
At the AAAI conference in Austin, Texas, in 1984, a panel of AI researchers warned that hopes for AI ran too high. Drew McDermott described a possible collapse of AI funding that some were calling an “AI winter”.
Why it matters It came true: interest and funding fell from the mid-1980s, and cheaper workstations undermined the companies making special AI computers, called Lisp machines.
Note This period is often called the second AI winter, but sources disagree on when it began: AI100 says the mid-1980s, others the late 1980s. The panel is the earliest use of the term we could confirm.
On 26 April 1985, the first ARM1 chip arrived at Acorn Computers from VLSI Technology and worked the same day. Sophie Wilson designed its instruction set, and Steve Furber its architecture.
Why it matters Its small, low-power design grew into the ARM architecture, spun out as a company in 1990 and now found in nearly every smartphone.
Note Some museum pages call ARM1 the first commercial RISC processor; we claim only that it was the first ARM chip.
In Nature, David Rumelhart, Geoffrey Hinton and Ronald Williams described backpropagation, a way to train neural networks with several layers by sending errors backwards. They showed that hidden layers can learn useful features of a task.
Why it matters It showed how to train networks with hidden layers, and backpropagation later trained deep networks such as LeCun’s.
Note The method is older: a survey traces it to Seppo Linnainmaa (1970) and its first use in neural networks to Paul Werbos (1981). The Nobel committee says the authors reinvented it; the 1986 paper made it popular.
Tim Berners-Lee, a scientist at CERN, wrote a proposal for a system of linked documents, so the lab would stop losing track of its information. Based on hypertext, it asked CERN’s managers to back the idea.
Why it matters It grew into the World Wide Web, which soon became the main way people used the internet.
Note The 1989 text did not yet use the name World Wide Web; Berners-Lee chose it in 1990. CERN and the W3C give only the month, March 1989.
Yann LeCun and colleagues at Bell Labs trained a neural network with backpropagation to read handwritten zip codes from US mail. Its design built in knowledge about images, so one network handled the whole job.
Why it matters Its design, later called a convolutional network, was used by several US banks to read handwritten cheques from the mid-1990s.
Note Fukushima’s neocognitron (1980) was a similar design. The new step was training such a network with backpropagation on real mail.
At IBM, Gerald Tesauro built a neural network that taught itself backgammon by playing against itself. In 1992, its second version played 38 games against top players and lost by only seven points in total.
Why it matters It showed that a program could learn to play near the level of the best humans from its own games.
Note Sources give different numbers of training games for the early version, so we give none.
Corinna Cortes and Vladimir Vapnik, at AT&T Bell Labs, described the support-vector network for sorting data into two groups. It worked even when the groups overlapped, and was tested on reading handwritten digits.
Why it matters It became a widely used method for analysing text, images and other data.
Note Support-vector ideas are older: Burges traces them to Vapnik’s work in the late 1970s, and the kernel idea was shown in 1992. The 1995 paper handled data that cannot be split cleanly.
On 11 May 1997, IBM’s Deep Blue won its six-game match against world chess champion Garry Kasparov in New York, 3.5 to 2.5. It is widely credited as the first computer to beat a reigning world champion in a match.
Why it matters It reached a goal pursued since the late 1940s and was widely seen as a symbolic contest between machine and human intelligence.
Note Kasparov accused the IBM team of cheating, IEEE Spectrum reports. Deep Blue had won one game against Kasparov in February 1996, but Kasparov won that match 4–2.
Sepp Hochreiter and Jürgen Schmidhuber described long short-term memory (LSTM), a recurrent neural network built from memory cells controlled by gates. In tests, it linked events more than 1,000 steps apart, which earlier recurrent networks struggled with.
Why it matters It let neural networks learn from long sequences and became an important method for speech and language.
Note The method was first described in a 1995 technical report, and Hochreiter’s 1991 thesis analysed the underlying problem. The date shown is the journal paper, November 1997.
In 1998, Sergey Brin and Larry Page of Stanford described Google, a search engine that ranks pages by the links pointing to them. Google Inc. was incorporated in California in September 1998.
Why it matters Ranking pages by their links, not only by keywords, became the foundation of Google’s search.
Note The paper is from the April 1998 issue. Google ties its birth to an investor’s cheque in August 1998, and its 2004 stock-market filing gives September 1998 for incorporation.
Nvidia launched the GeForce 256 graphics chip in 1999 and called it the world’s first GPU. It did the 3D shape and lighting calculations itself, which took work off the main processor.
Why it matters Nvidia made the name GPU famous, and GPUs later helped researchers train neural networks much faster.
Note “First GPU” is Nvidia’s own marketing claim: Jon Peddie notes that 3Dlabs and professional cards had similar engines earlier. Sources give August or October 1999, so only the year is shown.
On 20 November 2000, Honda showed ASIMO, a walking humanoid robot, at its Tokyo headquarters. At 120 cm tall, it was built to reach door handles and light switches in rooms made for people.
Why it matters It became one of the best-known humanoid robots, and Honda says its years of demonstrations taught lessons about robots working safely near people.
Note IEEE’s robot guide says Honda introduced ASIMO on 31 October 2000, the day Honda says it was created. The press release and public demonstration came on 20 November 2000.
In September 2002, iRobot began selling Roomba, a robot vacuum cleaner that cleans a floor on its own. It cost $199.95 and used sensors to follow walls and avoid stairs.
Why it matters Roomba brought robots into ordinary homes, and iRobot has reportedly sold about 40 million of them since.
Note iRobot’s press release is dated 18 September 2002; MIT Technology Review says Roomba was unveiled on 23 September. “The first automatic floor cleaner in the US” is iRobot’s own claim.
On 8 October 2005, Stanford’s driverless car Stanley won the DARPA Grand Challenge. It drove the 132-mile desert course in Nevada in under seven hours and won $2 million.
Why it matters Five cars finished after none had in 2004, and team leader Sebastian Thrun later worked on Google’s early self-driving cars.
Amazon Web Services launched S3, a service for renting online storage, on 14 March 2006. In August 2006 it opened a test version of EC2, which rents out computing power on demand.
Why it matters Teams could rent storage and computers and pay only for what they used, which started Amazon’s cloud business.
Note This milestone covers two launches. S3 launched on 14 March 2006, the date shown; the EC2 test version is dated 24 August 2006 by AWS and 25 August by TechCrunch.
In July 2006, Geoffrey Hinton, Simon Osindero and Yee-Whye Teh published a fast way to train neural networks with many layers. It trains one layer at a time, then fine-tunes the whole network.
Why it matters It made networks with many layers trainable and was a milestone on the road to what is now called deep learning.
Note Deep networks were studied before 2006: a survey by Jürgen Schmidhuber traces them back to 1965 and says the name “deep learning” took off around 2006.
On 9 January 2007, Steve Jobs showed the iPhone, which joined a phone, an iPod and an internet device in one touch-screen product. It went on sale in the US on 29 June 2007.
Why it matters It changed how people use phones, and Apple sold more than a billion iPhones within ten years.
Note The date is Apple’s announcement; sales began on 29 June 2007. The Computer History Museum notes that the iPhone was not the first smartphone.
In November 2006, Nvidia unveiled CUDA, a way for programmers to run general code on its graphics chips. A test version followed in February 2007, and CUDA 1.0 in June 2007.
Why it matters It opened GPUs to science and engineering work, and AlexNet’s creators used CUDA to train their 2012 image-recognition network.
Note Nvidia unveiled CUDA on 8 November 2006; a test version came on 15 February 2007 and CUDA 1.0 was announced on 26 June 2007, the date used here.
In June 2009, Fei-Fei Li and colleagues at Princeton presented ImageNet at the CVPR conference. It held 3.2 million labelled images in 5,247 categories, found online and checked by crowd workers.
Why it matters It supplied the data that AlexNet was trained on in 2012, a breakthrough for image recognition.
Note The Computer History Museum dates the project’s start to 2006. The database kept growing; ImageNet now lists more than 14 million images.
Demis Hassabis, Shane Legg and Mustafa Suleyman started DeepMind in London in 2010. Its stated aim was to “solve intelligence”, using ideas from machine learning and neuroscience.
Why it matters Google bought DeepMind in 2014, and the lab went on to build AlphaGo and AlphaFold.
Note DeepMind says it started in 2010, and the UK company register shows it was registered in September 2010 and renamed DeepMind Technologies in November 2010. MIT Technology Review (2014) says 2011.
IBM’s Watson played the quiz show Jeopardy! against champions Ken Jennings and Brad Rutter. The episodes aired on 14–16 February 2011, and Watson won with $77,147, against $24,000 and $21,600.
Why it matters Watson showed that a computer could answer questions asked in everyday language, and IBM later sold the technology to businesses.
On 4 October 2011, Apple announced Siri, a voice assistant built into the new iPhone 4S. People could ask it to make calls, send messages, set reminders and find places.
Why it matters It put a voice assistant on millions of phones: Apple said it sold over four million iPhone 4S in the first three days.
Note Siri began as a separate company, a spin-off of SRI, and first came out as an iPhone app in February 2010. Apple bought the company in April 2010.
A deep neural network built by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton won the 2012 ImageNet image-recognition contest. It missed the right answer in its top five guesses 15.3% of the time, against 26.2% for the best rival.
Why it matters AlexNet showed that deep networks, GPUs and big data sets could beat older methods, and neural networks soon took over computer vision.
Note Contest entries closed on 30 September 2012 and the results came out in October 2012. The paper followed in December 2012 at the NIPS conference.
In 2013, Tomas Mikolov and colleagues at Google showed a fast way to learn a vector, a list of numbers, for every word. Trained on 1.6 billion words in under a day, the vectors captured links such as country and capital.
Why it matters Word vectors became easy to train on huge texts, and NeurIPS later said the work began a new era in language processing.
Note Word vectors were not new in 2013, as the paper itself says; word2vec made them fast to train. The first paper came in January 2013, the code later that year and the second paper in October.
Ian Goodfellow and colleagues at the Université de Montréal described a new way to train models that create data, such as images. Two networks compete: a generator makes fakes, and a discriminator tries to spot them.
Why it matters GANs set off a wave of research on realistic generated images, and also raised fears about convincing fakes.
In September 2014, Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio published a new design for translation by neural networks. While writing each word, the model looks back at the source words that matter most.
Why it matters Attention became a standard part of language models, and in 2017 the transformer was built on attention alone.
In December 2015, Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun at Microsoft Research described residual networks with up to 152 layers. A group of them won the 2015 ImageNet classification contest, with 3.57% error.
Why it matters Shortcut connections made it possible to train networks with far more layers, and the transformer later used them too.
Note First posted on 10 December 2015, the day the ImageNet 2015 results came out. It won the best paper award at the CVPR conference in June 2016.
OpenAI was announced on 11 December 2015 as a non-profit AI research company. Backers including Elon Musk and Sam Altman pledged $1 billion, and Ilya Sutskever became research director.
Why it matters It created a well-funded lab that promised to share its research openly; it later built the GPT models and ChatGPT.
Note Reports describe the $1 billion differently: TechCrunch says “contributed”, The Christian Science Monitor says “investment”.
The European Parliament and the Council adopted Regulation (EU) 2016/679, known as the GDPR, dated 27 April 2016. It replaced a 1995 EU directive and applies in all member states from 25 May 2018.
Why it matters It limited decisions made only by machines, set fines of up to 4% of worldwide turnover, and inspired laws in Brazil, India and California.
Note “Adopted” covers several steps: Parliament voted on 14 April 2016, the act is dated 27 April, it entered into force on 24 May 2016 and it applies from 25 May 2018. We use the act’s date.
At its I/O conference in May 2016, Google announced the Tensor Processing Unit (TPU), a chip it built for machine learning. Google said it had already run TPUs in its data centres for more than a year.
Why it matters Google’s own 2017 paper reported that the TPU ran trained neural networks 15 to 30 times faster than the chips it was compared with.
Note Google’s own blog page is dated 19 May 2016, but news reports of the I/O announcement are dated 18 May. We give the month only.
Eight researchers working at Google posted a paper that introduced the transformer, a neural network design built on attention alone. It set new best scores on two translation tests and needed far less training time.
Why it matters It became the base of later language models, including BERT and the GPT family.
OpenAI researchers trained a transformer on thousands of books, then fine-tuned it on each task. It beat the best earlier results on 9 of the 12 tests the authors studied.
Why it matters It started OpenAI’s GPT line: GPT-2 and GPT-3 kept its design and trained larger models on more text.
Google researchers introduced BERT, a language model that uses the words on both sides of a word to understand it. It set new best results on eleven language tests, and Google shared the code in November 2018.
Why it matters Open code let anyone build on it, and in October 2019 Google said it was bringing BERT to Search.
OpenAI introduced GPT-2, a language model with 1.5 billion parameters that writes coherent paragraphs. Citing fears of misuse, it first shared only a small version and released the full model in November 2019.
Why it matters It sparked a wide debate on when and how AI labs should release powerful models.
Note Experts disagreed: some saw the staged release as a sensible precaution, others called it a publicity stunt. OpenAI later said it had seen no strong evidence of misuse.
On 22 May 2019, the OECD Council adopted a Recommendation on Artificial Intelligence, known as the OECD AI Principles. The OECD calls it the first intergovernmental standard on AI.
Why it matters In June 2019 the G20 leaders welcomed AI principles drawn from it, and the OECD updated it in 2023 and 2024.
OpenAI researchers described GPT-3, a language model with 175 billion parameters, over 100 times more than GPT-2. Given only a few examples in the prompt, it did well on many language tests without extra training.
Why it matters OpenAI opened an API to developers in June 2020, and the paper won a NeurIPS 2020 best paper award.
Note The few-shot results come from OpenAI’s own tests. The paper lists tasks where GPT-3 does poorly, and MIT Technology Review stressed that its fluent text hides a lack of real understanding.
DeepMind’s AlphaFold 2 took part in CASP14, a contest to predict the 3D shapes of proteins, and scored far ahead of the other teams. The organisers said about two thirds of its predictions matched lab-quality results.
Why it matters DeepMind published the method in Nature in July 2021, and scientists said it could speed up biology and drug research.
Note DeepMind called it a solution to a 50-year-old problem. CASP’s organisers noted limits: single proteins, not complexes. We do not say it “solved” protein folding.
OpenAI showed DALL·E, a neural network that creates images from written captions, such as an armchair shaped like an avocado. Like GPT-3, it is a transformer, trained on 250 million pairs of images and text.
Why it matters It showed that a GPT-style model can turn words into pictures; DALL·E 2 followed in April 2022, using a different method.
On 29 June 2021, GitHub opened a limited preview of Copilot, an AI tool that suggests whole lines or functions of code. It ran on OpenAI Codex, a GPT model fine-tuned on public code.
Why it matters It showed that AI could suggest code as programmers type, and by 2023 such tools were changing how software gets made.
Note Training on public code is contested: programmers have sued over it, MIT Technology Review reported in 2023.
On 22 August 2022, Stability AI released Stable Diffusion, a model that turns text prompts into images. The code and model files were free to download, under a licence with rules against misuse.
Why it matters Anyone with a good graphics card could run and adapt a capable image model, which also sparked debate about artists’ work in training data.
Note This was the public release; researchers had received the model earlier. “Open” here means downloadable model files under a licence with use rules, not classic open source.
On 30 November 2022, OpenAI released ChatGPT, a chatbot that answers questions in a back-and-forth conversation. OpenAI called it a research preview; its model was fine-tuned from GPT-3.5 with human feedback.
Why it matters By outside estimates it reached 100 million users in about two months, and rivals such as Google rushed out their own chatbots.
Note The 100 million users figure is an outside estimate, not an OpenAI count.
On 14 March 2023, OpenAI announced GPT-4, a model that accepts text and images and writes text. OpenAI said it scored around the top 10% of test takers on a simulated bar exam.
Why it matters Microsoft said Bing Chat already used a version of it, and OpenAI’s choice to withhold its size and training data drew criticism.
Note OpenAI did not publish the model’s size or training data, so size figures seen online are guesses. The exam results come from OpenAI’s own tests.
On 1 November 2023, 28 countries and the EU signed the Bletchley Declaration at the UK’s AI Safety Summit. It says the most capable AI models could cause serious harm, for example in cyber security and biotechnology.
Why it matters It started a series of summits, with South Korea and France hosting the next ones, and an expert report on AI risks led by Yoshua Bengio.
Note New Zealand joined on 23 October 2024, so the official list now shows 29 countries and the EU.
In December 2023, negotiators for the European Parliament and the EU’s member states reached a political deal on the AI Act. It sets rules by level of risk, bans some uses and covers general-purpose AI.
Why it matters It cleared the way for the final law: Parliament adopted it in March 2024 and the member states in May 2024.
Note Talks ended late on Friday 8 December; the official press releases are dated 9 December, so only the month is given.
On 1 August 2024, the EU AI Act, Regulation 2024/1689, entered into force. It sorts AI systems by level of risk and bans some uses, such as “social scoring” of people.
Why it matters Its rules switch on in stages: some bans applied from February 2025 and rules for general-purpose AI models from August 2025.
Note Later EU changes in 2026 moved some deadlines for high-risk systems; the stages named here did not change.
On 12 September 2024, OpenAI released o1-preview and o1-mini, models that work through a long chain of thought before they answer. OpenAI said they learned this by large-scale reinforcement learning.
Why it matters OpenAI reported large gains on maths and coding tests, and four months later DeepSeek-R1 claimed to match it.
Note Whether these models really reason is debated. Asking models to write out their steps was shown earlier, in a January 2022 paper, so we don’t call o1 the first.
The 2024 Nobel Prize in Physics went to John Hopfield and Geoffrey Hinton for work on artificial neural networks. The Chemistry prize went to David Baker for protein design, and to Demis Hassabis and John Jumper for protein structure prediction.
Why it matters Two science Nobels in two days honoured AI work; the Academy said AlphaFold2 had predicted about 200 million protein structures.
Note The two prizes were announced on different days, Physics on 8 October and Chemistry on 9 October 2024, so only the month is given.
On 20 January 2025, the Chinese company DeepSeek released DeepSeek-R1 with open weights under an MIT licence. DeepSeek said it matched OpenAI’s o1 on maths, code and reasoning tasks.
Why it matters It put a strong reasoning model in everyone’s hands for free, and rivals such as Alibaba and OpenAI reacted quickly.
Note Not the first open reasoning model: IEEE Spectrum names Alibaba’s QwQ as an earlier one. A web-only preview, R1-Lite, came on 20 November 2024.
A milestone goes in when it changed what AI can do, the tools AI is built with, how many people use AI, or the rules for AI. It must be at least a year old: newer events are news.
We checked every milestone in at least two sources and read them: the original paper, patent, announcement or law where it exists, and a history source such as a museum, an encyclopedia or a report from the time. We never cite Wikipedia. Many “firsts” are disputed, so we say when they are, and dates are only as exact as the sources agree.
Last checked 4 October 2026. Found a mistake? Tell us, and we will correct it. How we report