Recent history: why it worked now and not before

The central ideas of deep learning are old. Artificial neural networks date from the middle of the twentieth century, the algorithm that trains them became popular in the eighties, and the field still spent decades promising things it could not deliver. The interesting question is not where the ideas came from: it is why they only started working now.

The short answer is that the theory did not change, the scale did. v1 of this guide counted three factors, all three turning on the same concept —the datum—: producing data, storing it and processing it. Two more are needed today.

Processing: machines that do more per second

A convenient measure of progress is the number of transistors on a processor. In 1965 Gordon Moore observed that the number had been doubling every year since the transistor was invented, and ventured that it would keep doing so for another decade. The observation became known as “Moore’s law”, though there is little law about it: it is an industrial trend, and one running out against the physical limits of miniaturisation.

What matters is the shape of the curve. If capacity doubles every two years, then in four years there is not twice as much: there is far more than twice as much. The 1971 Intel 4004 had 2,300 transistors on 12 × 12 millimetres; current processors have billions on less area.

With one correction that matters: the hardware that made deep learning possible is not the general-purpose processor but the graphics card, designed for video games and repurposed because rendering polygons and multiplying matrices turn out to be the same problem. Much of today’s semiconductor geopolitics is explained by that accident.

Production: data everywhere

Almost everything we do can be turned into data, because there is a digital mediation that registers the smallest change of state:

  • Clicks on a page generate data about which content draws attention.
  • Searches generate data about what people want to know.
  • Likes, comments and shares generate data about popularity.
  • Histories and cookies generate data about browsing patterns.
  • Product reviews generate data about consumer preferences.
  • The phone’s GPS generates data about people’s movements.
  • Connected devices generate environmental, fitness, energy-consumption and industrial-production data.

Luciano Floridi called the infosphere the informational environment that digital technologies created around the planet: the totality of what is produced, stored, shared and processed by digital means. His thesis is that just as living organisms interact with the biosphere and constitute it, we interact with the infosphere and constitute it. Hence his proposal that we think of ourselves as informational organisms living an onlife existence, with no sharp border between connected and unconnected.

The metaphor has a limit worth keeping in mind, and it is the same one “the cloud” has: it names as ethereal something that sits in a shed, consumes electricity and occupies a plot with a postal address.

Storage: having somewhere to put it

There is so much data available partly because there is somewhere to keep it, and that is as hard a technological problem as processing. Here the decisive factor is economic: the cost per gigabyte fell so many orders of magnitude that it stopped being a design constraint. In the late eighties, storing a gigabyte cost on the order of a hundred thousand dollars; today it is measured in cents.

Put on Turing’s scale: the gigabyte he estimated the imitation game would need would have cost, on 1950s disks, on the order of nine million dollars, and in fast memory a figure not worth writing down because it was also technically impossible.

Source: Our World in Data.

Notice anything odd about the shape of the chart? Look at the scale on the vertical axis.

Architecture: the transformer

The three factors above explain the possibility, not the result. The missing piece arrived in 2017 with the transformer, the architecture proposed in Attention is all you need, whose decisive advantage is not so much that it understands language better as that it parallelises: it can use thousands of cards at once, where earlier architectures processed in sequence.

That turned the problem of artificial intelligence into a budget problem. If more compute and more data systematically produce better models, the question stops being “what idea are we missing?” and becomes “who can pay for the next training run?”. The public break came in late 2022, when that technical shift became a mass consumer product.

And what the scale costs

This is where v1 of this guide stopped, and today it cannot. Scale is not free: training a large model consumes electricity and water, happens in data centres sited where energy and land are cheap —which rarely coincides with where the people using the model live— and depends on badly paid human labour to label data and correct outputs.

This is not an ethical appendix to the technical matter: it is part of the explanation of why this worked now. A model that is only viable because cooling water and labelling work are cheap is an artefact whose condition of possibility is an unequal distribution. That argument, with its sources, is in Thinking AI from the Global South.

So what is a neural network

Let us join the dots. When people talk about “the algorithm” of Netflix, Spotify or a social network, strictly speaking they are not talking about an algorithm: an algorithm guarantees a result, and there is no such guarantee here. What there is, is a set of computational processes —those are algorithms, and fairly simple ones to program— that configure an artificial neural network able to perform a task from the state it reached by exploring data.

A neural network is a layered logical structure, with nodes that receive signals, process them and pass the result to other nodes. What is distinctive are the weighted connections between nodes, the weights. Training consists of iteratively adjusting those weights according to feedback from the data, until the network models the relation between what goes in and what we want to come out. For a great many tasks you do not need precise instructions on how to tell a dog from a cat: plenty of examples and plenty of compute will do.

And that is where the problem that opens the next entry appears. We solved one thing and created another: if I do not know exactly what the network is doing to achieve what I asked of it, how do I know it is doing it well, and that it will keep doing it well faced with something it has never seen? Put differently: how do I know that it knows? And, before that, how do I know that I know anything?

Going deeper

Floridi, L. (2014). The fourth revolution: How the infosphere is reshaping human reality. Oxford University Press.

Ilcic, A. A. Desafíos epistémicos de la técnica: ¿aprenden las máquinas con su machine learning? Manuscript, in Spanish.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł. and Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. https://arxiv.org/abs/1706.03762 Open access

If you are not afraid of the mathematics, the full technical explanation is in the route laid out in Getting started, and in particular in the 3Blue1Brown series and Karpathy’s Zero to Hero.

Videos on the history of AI

The machine that changed the world

The Thinking Machine (MIT, 1961), included in this compilation of documentaries on the history of AI:

This entry revises and updates Recent history from v1, available in Spanish.

docs