AI news, models and products, with sources中文
NewsLandmark Research15 Mar 2016

AlphaGo beats Lee Sedol 4–1, years ahead of expert predictions

DeepMind's AlphaGo beat 18-time world champion Lee Sedol 4–1 in Seoul in a game with about 10^170 possible positions, where most experts had expected computers to need at least five to ten more years. Its Move 37, which it estimated a human would play 1 time in 10,000, showed a machine could make creative judgements, not just calculate.

What happened

Go player Lee Sedol smiling
Lee Sedol in May 2016, two months after the match. Photo: LG Electronics / Wikimedia Commons, CC BY 2.0

From 9 to 15 March 2016, at the Four Seasons Hotel in Seoul. On one side, Lee Sedol: a 9-dan professional with 18 international titles, regarded as one of the strongest players in the history of Go. On the other, AlphaGo, a program built by Google’s DeepMind, with team member Aja Huang placing its stones on the board.

Five games, a $1 million prize. AlphaGo won games 1, 2, 3 and 5; Lee won game 4. Every game ended by resignation.

Final score (AlphaGo–Lee)
4 : 1
People who watched (DeepMind)
200m+
Possible Go positions
10¹⁷⁰
Chance a human plays Move 37
1 in 10,000

Sources: DeepMind's AlphaGo page; Wikipedia, “AlphaGo versus Lee Sedol”

Game Date Black Result Highlight
1 9 March Lee AlphaGo wins Lee looked in control until the last 20 minutes
2 10 March AlphaGo AlphaGo wins AlphaGo’s surprising Move 37; Lee thinks for a long time
3 12 March Lee AlphaGo wins Analysts called AlphaGo’s strength “almost scary”
4 13 March AlphaGo Lee wins Lee’s Move 78 turns the game — humanity’s only win
5 15 March Lee AlphaGo wins A close game decided in the endgame

Afterwards the Korea Baduk Association awarded AlphaGo an honorary 9-dan rank. DeepMind gave the $1 million prize to charities including UNICEF and to Go organisations; Lee received $170,000 for playing and for his win.

How big was it? Why “a decade early”

Difficulty: you can’t just calculate

When IBM’s Deep Blue beat world chess champion Garry Kasparov in 1997 (see Deep Blue beats the world chess champion), it relied mainly on raw calculation: search an enormous number of moves and pick the best. That doesn’t work for Go.

A Go board has 19×19 = 361 points; at the start every move has hundreds of options, and a game usually runs to one or two hundred moves. DeepMind puts the number of possible positions at about 10 to the power 170 — more than the atoms in the observable universe, and “a googol times” more complex than chess. No computer could ever search them all.

Judging a position is even harder. In chess you can count material and check the king’s safety; a Go master’s judgement is often an intuition they can’t fully explain. As the mathematician I. J. Good wrote in 1965, programming a computer to play a reasonable game of Go would be harder than chess, because the principles are “more qualitative and mysterious … and depend more on judgement”.

Before 2015 the best Go programs played at amateur level and needed a four- or five-stone handicap to beat a 9-dan professional. Most observers expected Lee to win; according to Wikipedia’s summary, most experts thought a program as strong as AlphaGo was at least five years away, some said ten. Lee himself predicted a “landslide”.

Comparison: how fast AlphaGo itself improved

Playing strength is measured with Elo ratings: a 200-point gap means the stronger side wins about 75% of the time; 400 points, about 90%. DeepMind’s 2017 Nature paper gave ratings for each version:

Elo ratings of AlphaGo versionsHigher is stronger; results in brackets
  • AlphaGo Fan (Oct 2015, 5–0 v Fan Hui)3,144
  • AlphaGo Lee (Mar 2016, 4–1 v Lee Sedol)3,739
  • AlphaGo Master (2017, 60–0 online)4,858
  • AlphaGo Zero (Oct 2017, 100–0 v Lee version)5,185

Ratings were estimated by DeepMind from games between versions and are not directly comparable to human rating lists. Source: Silver et al., Nature 2017 (AlphaGo Zero paper), as compiled by Wikipedia's “AlphaGo” article

Only five months separated beating European champion Fan Hui (a 2-dan professional) in October 2015 and beating Lee Sedol in March 2016. And the version that beat Lee was itself beaten 100–0 eighteen months later by AlphaGo Zero, which never studied a human game.

Move 37: a move no human would play

In game 2, with AlphaGo playing black, Move 37 landed on the fifth line on the right side — a “shoulder hit” almost no professional would consider. The English commentator, US 9-dan professional Michael Redmond, called it “creative” and “unique”. Lee took an unusually long time to reply.

DeepMind later said that, by AlphaGo’s own estimate learned from human games, the chance of a human playing that move was 1 in 10,000. It played it anyway, because its own judgement said it improved its chances. AlphaGo won the game; Lee said afterwards that “AlphaGo played a nearly perfect game”.

Here is the full game record. Step through with ▶, or jump straight to Move 37:

Game 2 (10 March 2016): AlphaGo plays black● AlphaGo ○ Lee Sedol

Game record: Google DeepMind Challenge Match, game 2; SGF from the CWI Go games archive. Board drawn by YLEM from the record.

Move 78: humanity’s only win

After three losses Lee changed strategy in game 4: take the corners and edges, then invade AlphaGo’s sphere of influence to create an all-or-nothing fight. With Move 78 he played a wedge in the centre that the Chinese professional Gu Li called a “divine move”. DeepMind later said this move, too, had about a 1 in 10,000 chance of being played.

AlphaGo’s reply went wrong: at Move 79 it still estimated a 70% chance of winning; only by Move 87 did it realise things had collapsed, and it then played a run of clearly bad moves. Lee won, calling it a “priceless win that I would not exchange for anything”.

Game 4 (13 March 2016): Lee Sedol plays white● AlphaGo ○ Lee Sedol

Game record: Google DeepMind Challenge Match, game 4; SGF from the CWI Go games archive. Board drawn by YLEM from the record.

How AlphaGo did it

Nobody wrote a Go rulebook of good moves into AlphaGo; its judgement was learned. Roughly:

How AlphaGo learned to play
  1. 1Imitate humansStudy about 30 million moves from 160,000 strong amateur games to learn where good players usually play.
  2. 2Play itselfPlay tens of millions of games against versions of itself; strengthen what wins (reinforcement learning).
  3. 3Two intuition networksOne suggests which few moves are worth considering; the other estimates who is winning.
  4. 4Search selectivelyRead ahead only along the most promising lines (Monte Carlo tree search) instead of everything.

Sources: Wikipedia, “AlphaGo” and “AlphaGo versus Lee Sedol”, summarising DeepMind's 2016 Nature paper

The key is the third step: AlphaGo had something like a player’s intuition, ruling out most bad moves at a glance and spending its calculation on a few lines. Against Lee it ran on Google’s cloud, using many processors and Google’s own TPU chips.

One trait puzzled viewers: it maximises its probability of winning, not its margin. Offered an 80% chance of winning by 20 points or a 99% chance of winning by 1.5, it takes the second. So when ahead it often played moves that looked “slack” but were simply safe.

What happened next

From Seoul to today
  1. 1Early 2017As "Master", it wins 60–0 against top players in online fast games.
  2. 2May 2017Beats Ke Jie, then the world's top-ranked player, 3–0.
  3. 3Oct 2017AlphaGo Zero, trained with no human games, beats the Lee version 100–0 after three days.
  4. 4Nov 2019Lee retires: "Even if I become the number one, there is an entity that cannot be defeated."

Sources: Wikipedia, “AlphaGo”, “AlphaGo Zero”, “Lee Sedol”

For AI research: learning intuition with neural networks, combined with search and self-play, was extended by DeepMind to chess and shogi (AlphaZero); the team’s deep-learning experience then went into protein-structure prediction (AlphaFold 2). DeepMind’s chief executive Demis Hassabis later shared the 2024 Nobel Prize in Chemistry for AlphaFold (see Two Nobel Prizes for AI research). Today’s “reasoning” language models also rely heavily on reinforcement learning (see o1).

For Go itself: a study published in the Proceedings of the National Academy of Sciences (PNAS) in 2023 analysed more than 5.8 million moves by professional players from 1950 to 2021, scoring each with a superhuman Go AI. It found that human decision quality improved significantly after superhuman AI appeared, and that players made more never-before-seen moves, which were also of higher quality. AlphaGo didn’t end Go; it freed players from centuries of convention.

For Lee Sedol: he retired in 2019. In a 2024 interview he said: “Losing to AI, in a sense, meant my entire world was collapsing … I could no longer enjoy the game.”

For society: two days after the match, South Korea announced it would invest 1 trillion won (about $863 million) in AI research over five years. Deep Blue’s Murray Campbell called it “the end of an era … board games are more or less done and it’s time to move on.”

Things to keep in mind

  • It only plays Go. AlphaGo was trained for one game. It can’t chat or write and is a different kind of system from ChatGPT or Claude. It showed the power of a method, not general intelligence.
  • It had weaknesses. Game 4 showed AlphaGo could misjudge an unexpected move; commentators linked this to how Monte Carlo search “prunes” lines that look unimportant.
  • Unequal resources. One human brain against a lot of chips in Google’s data centres. That doesn’t change the result, but the match was also a show of engineering muscle.
  • Some numbers come from DeepMind. The 200 million viewers, the 1-in-10,000 odds and the version ratings were published by DeepMind; the results and game records are public facts.

What it means for us today

For many people AlphaGo was the first moment they realised AI might surpass humans at something that takes intuition. Ten years on, the same thing is happening in programming, mathematics and science: in 2025 AI reached gold-medal level at the International Mathematical Olympiad (see AI reaches gold-medal level at the Maths Olympiad), and in 2026 AI began claiming proofs of open research problems (see OpenAI’s 700+ AI-written maths papers).

Go’s experience may be the most useful lesson: after AI overtook humans, the game didn’t disappear — players got better by using AI as a teacher and sparring partner, not only as an opponent.

Watch the documentary

In 2020 DeepMind made its award-winning documentary AlphaGo free on its official YouTube channel. It follows all five games in Seoul, including the moments around Move 37 and Move 78.

AlphaGo (2017), the award-winning documentary. Open on YouTubeYouTube loads only after you click.

Sources