AI Data Center (AI 데이터 센터)
Overview
An AI data center (AI 데이터 센터) is a specialized computing facility designed to handle the training and inference workloads of large language models (LLMs) and generative AI. Whereas conventional data centers were optimized for CPU-centric web services and database processing, AI data centers are characterized by the high-density integration of thousands to tens of thousands of accelerators such as GPUs and TPUs, combined with ultra-high-speed interconnects and large-capacity power and cooling infrastructure. Since 2023, alongside the generative AI boom, the industry's center of gravity has been shifting from "cloud data centers" to "AI factories," and the structure is being reorganized so that securing power and land is itself a source of competitiveness.
Key Details
Definition and Distinctions
AI data centers are broadly divided into two types according to purpose. The first is the training cluster, which connects tens of thousands of accelerators via NVLink and InfiniBand to carry out large-scale computation over several months. The second is the inference cluster, which focuses on optimizing latency, power efficiency, and tokens-per-second throughput. Recently, hybrid configurations that handle both training and inference within a single facility have been increasing.
Hardware Composition
- Accelerators: NVIDIA H100/H200, Blackwell B200/GB200, AMD MI300X/MI325X, Google TPU (v5e/v5p/Trillium), and AWS Trainium/Inferentia are representative.
- Memory: HBM3 and HBM3E are the key variables for performance, and HBM supply volume dictates overall system shipments.
- Server form factor: Moving beyond the HGX approach of binding eight GPUs onto a single baseboard, the NVL72 form, which binds 72 GPUs into a single domain at the rack level, has emerged as the mainstream.
- Storage: NVMe SSDs and parallel file systems such as Lustre and Weka handle petabyte-scale training data.
Power and Cooling Infrastructure
The biggest bottleneck for AI data centers is power. While traditional rack power density was 5–15 kW, AI racks exceed 40 kW and reach 100–120 kW. Accordingly, air cooling has reached its limits, and direct-to-chip (D2C) cooling and immersion cooling are spreading rapidly. On the power side, competition is unfolding over long-term power purchase agreements (PPAs), gas turbines, and securing small modular reactor (SMR) and existing nuclear power. PUE (Power Usage Effectiveness) and WUE (Water Usage Effectiveness) are used as efficiency metrics.
Network and Software
The larger the cluster, the more the network determines performance. InfiniBand NDR/XDR, NVIDIA Spectrum-X Ethernet, and RoCEv2 are mainly used, and demand for 800G and 1.6T optical modules has surged. The software stack consists of the CUDA and ROCm ecosystems, Kubernetes/Slurm/Ray-based orchestration, and distributed training frameworks that combine data, tensor, and pipeline parallelism.
Operators and Market Structure
Hyperscalers (Microsoft Azure, AWS, Google Cloud, Meta, Oracle OCI) drive captive demand, while neoclouds such as CoreWeave and Lambda Labs provide specialized services. Colocation operators (Equinix, Digital Realty) supply power and land, and the semiconductor value chain (NVIDIA, AMD, Broadcom, TSMC, SK Hynix) forms the upstream of the entire ecosystem. In Korea, Naver, KT, SK Telecom, Samsung SDS, and LG are expanding infrastructure in connection with the government's GPU procurement projects.
Economics and Challenges
AI data centers are a massively capital-intensive business. GPU depreciation cycles (typically 3–5 years), a rising share of power costs, and $/GPU-hour price competition determine profitability. Major challenges include ① grid connection queues and transmission and distribution bottlenecks, ② land, permitting, and community acceptance, ③ cooling water use and carbon emissions, ④ HBM and advanced packaging supply chain constraints, and ⑤ debate over a bubble in investment returns.
Latest Trends
The biggest change in 2024–2025 is the emergence of gigawatt (GW)-scale campuses. Representative examples include the Stargate project pursued by OpenAI, Oracle, and SoftProject, and xAI's Memphis Colossus (100,000–1,000,000 GPU scale). NVIDIA is shipping GB200 NVL72 as a rack-level finished product, leading the standardization of liquid cooling, and discussions on standardizing rack, power, and cooling specifications are active centered on OCP (Open Compute Project).
The competition to secure power has also entered a new phase. Microsoft signed a power purchase agreement for the restart of the Three Mile Island nuclear plant, Google invested in Kairos Power's SMR, and Amazon signed a nuclear power contract with Talen Energy. Matching data center power consumption with 24/7 carbon-free energy (CFE) has become a key metric.
Technologically, the main trends are ① the rapid growth of the inference market and distribution toward small models and on-device AI, ② the adoption of optical interconnects and silicon photonics, ③ chiplets and 3D packaging that improve cooling efficiency, and ④ the expansion of China's own chips (Huawei Ascend) and AI infrastructure investment in the Middle East (Saudi Arabia, UAE) amid geopolitical export controls. On the regulatory side, the EU AI Act and mandatory energy and environmental reporting are directly affecting data center site selection.
Related Topics
- [[GPU]]
- [[Generative AI]]
- [[NVIDIA]]
- [[Cloud Computing]]
- [[Liquid Cooling]]
- [[Data Center]]