How AI emerged from its winters
AI's current rise rests on decades of learning, data and computing rather than a sudden invention. Its two earlier winters show why executives should pair ambition with evidence.
AI’s current prominence can look sudden. It is better understood as the result of seven decades of uneven progress. That longer view explains both why the present wave has substance and why confidence should still be tempered by evidence.
Ambition, rules and retreat
The modern field took shape in the 1950s, when researchers began asking whether machines could perform tasks associated with human intelligence. Expectations rose quickly. Some early predictions suggested that broad machine intelligence was close at hand.
In the summer of 1956, the Dartmouth Summer Research Project on Artificial Intelligence convened for two months at Dartmouth College. The name “artificial intelligence” first appeared in the 1955 proposal written by John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon. Participants including Allen Newell and Herbert Simon brought many of the era’s leading thinkers into the same room.
Every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it.

Proposal for the Dartmouth Summer Research Project on Artificial Intelligence, 31 August 1955. Signed by J. McCarthy, M. L. Minsky, N. Rochester and C. E. Shannon.
The proposal expected that a team of ten could make meaningful progress in two months. In retrospect the timetable is as striking as the ambition. Seventy years later the field is still working on some of the same questions, and the pattern of high expectations followed by disappointment was already taking shape.
The first systems largely depended on rules written by people. An expert would describe a problem step by step, and an engineer would translate that knowledge into instructions the computer could follow. This worked in narrow, orderly settings. It struggled when the real world produced ambiguity, exceptions and more combinations than anyone could anticipate.
The mismatch between ambition and performance led to the first AI winter in the 1970s. Funding and public attention fell sharply as promised capabilities failed to arrive. Progress continued in parts of the field, but the sense of an imminent breakthrough disappeared.
Optimism returned in the 1980s around expert systems, which encoded specialist judgement as large collections of rules. They found useful applications, yet were costly to build, difficult to maintain and brittle when conditions changed. A second winter began in the late 1980s and extended into the 1990s. Investment and interest collapsed again.
The two winters were not proof that intelligent systems were impossible. They showed what happens when claims move much faster than practical results.
The break from rules to learning
The decisive break came from changing how the problem was approached. Instead of trying to describe every rule, researchers increasingly built systems that learned patterns from examples. Deep learning, a method that learns through many layers of pattern recognition, took this further. The system could work out which features mattered rather than waiting for a person to specify them.
Consider image recognition. A rule-based approach might tell a computer that a cat has whiskers, triangular ears and a particular outline. Those descriptions soon fail across different breeds, poses, lighting and backgrounds. A learning system is shown many labelled images and finds the distinguishing features for itself.
The method was essential, but it was not sufficient. Three enabling conditions matured together. Digitisation created vast stores of images, text and other data. Processing power kept rising over decades, a trend known as Moore’s law. Graphics processing units, or GPUs, were developed to render computer games but proved unusually well suited to deep learning. They can repeat the same operation thousands of times at once, which makes training far faster.
Around 2012, deep learning beat carefully hand-designed approaches in a major image-recognition competition by a wide margin. The result was a practical turning point. It showed that learning from large amounts of data, supported by specialised computing, could outperform methods built around human-selected features. Advances in speech, language and vision followed.
Today’s large language models continue the same underlying idea. They learn patterns from examples rather than receiving a complete set of written rules. What has changed most is the scale of the data and computing, alongside far more disciplined engineering.
For executives, this history supports neither dismissal nor uncritical enthusiasm. The current wave rests on long accumulation, not a passing fashion. Yet excessive expectations stalled the field twice. The useful response is to test claims against work that matters to the organization, measure performance against a clear baseline, and expand only where the evidence holds.