AGI Criteria

A document that summarizes discussions on the definition and evaluation criteria of AGI (Artificial General Intelligence) and the latest trends from 2024 to 2025.

AGI Criteria

Overview

AGI (Artificial General Intelligence) refers to an intelligent system that can exhibit human-like learning, reasoning, planning, and creativity without being limited to specific tasks. However, there is still no consensus in academia and industry on exactly what AGI is or by what criteria its attainment should be judged. This document addresses the various criteria and classification systems for defining and evaluating AGI, along with recent technical debates.

Main Content

Definition of AGI and Historical Background

The concept of AGI has existed since the early days of computer science, and the Turing Test proposed by Alan Turing was an early criterion for determining whether a machine could converse at a human-equivalent level. Later, artificial intelligence research became concentrated on narrow AI, which solves specific problems, while AGI remained a long-term goal. With recent advances in deep learning and the emergence of large language models (LLMs), discussions on the feasibility of building AGI have become more active.

Main Criteria for Judging AGI

Human-level Generality and Adaptability

AGI must possess the generality to perform a wide variety of tasks and the ability to adapt to new environments. In other words, rather than superhuman performance in a single domain, the ability to generalize to tasks that have not been learned is considered a core criterion. This includes commonsense reasoning, causal understanding, transfer learning, and the ability to set goals and plan independently.

Definitions Centered on Performance and Economic Value

OpenAI defines AGI in its charter as 'highly autonomous systems that outperform humans at most economically valuable work.' This definition focuses less on technical completeness and more on the ability to replace actual human labor and produce value. It is a practical standard from an industrial perspective, but it has also been criticized for reducing the concept of intelligence too heavily to economic output.

DeepMind's 'Levels of AGI' Classification

In the paper 'Levels of AGI' published by Google DeepMind researchers in 2023, AGI is divided into six levels (Level 0 to 5) based on the two axes of performance and generality. Level 0 is narrow AI that does not reach human ability; Level 1 is 'Emerging,' performing specific tasks at a human-equivalent level; Level 2 is 'Competent,' demonstrating human-level performance in a wide range of tasks; Level 3 is 'Expert,' performing at human expert levels in multiple domains; Level 4 is 'Virtuoso,' surpassing humans in most fields; and Level 5 is 'Superhuman,' surpassing humans in all fields. This classification has been evaluated as presenting a clear standard that can measure not only whether AGI has been achieved but also gradual progress.

Limitations of Evaluation Benchmarks and Alternatives

Problems with Existing Task-Specific Benchmarks

Benchmarks such as GLUE, SuperGLUE, MMLU, and HumanEval evaluate particular NLP problems or coding abilities, but they are merely tools for measuring narrow AI performance and are difficult to regard as evidence of AGI attainment. There are concerns that LLMs repeatedly learn from similar problems and become over-optimized for benchmarks, distorting the measurement of genuine intelligence.

Emergence of Evaluations Dedicated to AGI

In response, new evaluation methods such as ARC-AGI (Abstraction and Reasoning Corpus) have been proposed. These consist of problems that test analogy and abstraction without relying on pretrained knowledge, and LLMs still have low accuracy on them, suggesting that AGI remains far off. Meanwhile, OpenAI's 'o1' model far surpassed existing LLMs on math and science problems requiring complex reasoning, demonstrating the potential of 'reasoning AI,' but it also revealed limitations in generalization ability.

Philosophical Debates on Judging AGI Attainment

What Does Human-Level Ability Mean?

Some scholars set the standard of AGI as a level indistinguishable from human behavior, while others argue that the expression 'human level' itself is ambiguous when considering superhuman processing speed and data scale that human intelligence never possesses. For example, LLMs are already better than most humans at multilingual processing, but in everyday physical-world conversation and interaction, they fall short of young children.

Subjective Indicators and Cultural Differences

The definition of intelligence may vary depending on culture and values. Some societies may emphasize linguistic ability, while others may emphasize negotiation skills or emotional empathy. Therefore, attempts to define AGI with a single objective indicator may involve political and ethical judgments rather than being purely scientific.

Connection to Industry and Safety Discussions

AGI criteria directly influence not only technical evaluation issues but also the design of AI safety and regulatory policies. Major research institutions, including OpenAI, DeepMind, and Anthropic, have put forward their own 'safety criteria' to prevent the development of AGI beyond human control. In particular, since 2024, the term 'AGI readiness' has emerged, along with moves to establish staged regulations according to economic impact.

Recent Trends (2024-2025)

Since 2024, the center of gravity in AGI discussions has shifted from 'capability' to 'directionality' and 'controllability.' OpenAI has stated that its large-scale models have reached a 'near-AGI level' by its internal standards, but that it is adjusting the stages of release due to safety concerns. The six-level classification proposed by Google DeepMind is gradually becoming accepted as an industry standard, and many AI companies are introducing their models according to those levels. In early 2025, some models recorded accuracy above 50% on the ARC-AGI test, leading to assessments that they had moved one step closer to AGI, yet the prevailing expert opinion is that they are still far from human-level general intelligence. Also, regulatory frameworks such as the EU's AI Act have begun classifying AGI as a separate category, and there is a growing voice advocating 'IA (artificial intelligence, a system that collaborates with humans)' rather than the development of superhuman intelligence.

Related Topics

  • [[인공 일반 지능]]
  • [[튜링 테스트]]
  • [[AI 안전]]
  • [[OpenAI]]
  • [[DeepMind]]
  • [[ARC-AGI]]