AlphaGo
Overview
AlphaGo is an artificial intelligence Go program developed by Google DeepMind. By combining deep learning neural networks and Monte Carlo tree search (MCTS), it surpassed the limitations of existing computer Go programs, and by winning its match against 9-dan Lee Sedol in 2016, it simultaneously showed the world both the potential and the shock of artificial intelligence. This match is evaluated as a signal that AI could encroach even on intuition and creativity, which had been considered uniquely human domains, beyond the simple win or loss of a game.
Key Details
Development Background
Go is a representative 'hard problem' game known to have more possible positions than the number of atoms in the universe. Even after chess was conquered by IBM's Deep Blue in 1997, Go remained for nearly 20 years the last bastion where the best human players overwhelmed computers. DeepMind launched the AlphaGo project in earnest in 2014 to tackle this problem head-on.
Core Technology
AlphaGo's core consists of three main elements.
- Policy Network: A neural network that predicts the next move, dramatically reducing the search space.
- Value Network: A neural network that evaluates the win rate of the current position, judging the advantage or disadvantage of a position without simulating to the end.
- Monte Carlo Tree Search (MCTS): Combines the above two neural networks into the search process to find the optimal move.
To this were added reinforcement learning through self-play and supervised learning using professional game record data.
Major Matches
- October 2015: Announced its existence by defeating European champion Fan Hui 2-dan in a five-game sweep.
- March 2016: Won 4-1 in a five-game match against 9-dan Lee Sedol. Lee Sedol's 'move 78' is remembered as the only victory a human achieved against AlphaGo.
- May 2017: In Wuzhen (乌镇), China, defeated 9-dan Ke Jie 3-0, achieving a complete victory against the world No. 1.
AlphaGo Zero and Afterward
In 2017, DeepMind announced AlphaGo Zero, which learned only through self-play without human game records, and AlphaZero, a general-purpose algorithm applicable to games other than Go. AlphaGo Zero overwhelmed the existing AlphaGo 100:0, proving that learning is possible without human knowledge.
Latest Trends
AlphaGo itself retired after 2017, but its technological legacy has deeply permeated the AI industry as a whole in 2024–2025.
- AlphaFold series: DeepMind's AlphaFold solved the protein structure prediction problem and made a decisive contribution to winning the 2024 Nobel Prize in Chemistry.
- Revival of Reinforcement Learning: AlphaGo's self-play method is being reexamined as the prototype of modern generative AI training techniques, such as reinforcement of reasoning in large language models (LLMs) and RLHF (reinforcement learning from human feedback).
- South Korea's AI Ecosystem: With the Lee Sedol match as a catalyst, South Korea began in earnest to foster AI talent and invest in policy, and since 2024, discussions on the 'AI Basic Act' and strengthening generative AI competitiveness have emerged as major tasks.
- Reappraisal: In 2024, Lee Sedol's 'move 78' was reanalyzed in AI research as an example of original strategy, and the AlphaGo match is still used as educational and research content.
Related Topics
- [[DeepMind]]
- [[Lee Sedol]]
- [[Reinforcement Learning]]
- [[AlphaFold]]
- [[Artificial Intelligence]]