Every System · Tracked
Every rack-scale AI system from every manufacturer. GPU counts, FP8 PFLOPS, power draw, cooling type, TCO per PFLOPS, and known datacenter deployments · sourced from spec sheets, earnings calls, and disclosed infrastructure builds.
- $973
- $1,469
- $1,768
- $2,063
- $3,306
- $14,734
- DGX GB300 NVL72 offers the lowest TCO at $973/PFLOPS/year
- Liquid-cooled systems average 6% lower TCO than air-cooled
- 4 of 8 ranked systems are NVIDIA-based
Systems table
10 systems · sorted by FP8 PFLOPS · all manufacturers
| System | GPUs | FP8 PFLOPS |
|---|---|---|
TPU v5p Pod Google | 8,960 | 8,100 |
DGX GB300 NVL72 NVIDIA | 72 | 1,440 |
DGX GB200 NVL72 NVIDIA | 72 | 720 |
TPU v6e Pod Google | 256 | 230 |
Maia 100 Rack Microsoft | 32 | 96 |
DGX B200 NVIDIA | 8 | 80 |
MI325X Platform AMD | 8 | 48 |
Trn2 UltraServer AWS | 16 | 48 |
HGX H100 NVIDIA | 8 | 32 |
CS-3 Cerebras | 1 | n/a |
TCO comparison
$/PFLOPS/year · hardware amortized 3yr + power at $0.05/kWh · lower is better
Cooling breakdown
Liquid vs air · power stats · efficiency comparison
Liquid cooling enables PUE of 1.05 to 1.15 vs 1.3 to 1.5 for air cooling. At datacenter scale, this translates to 15 to 30% lower power costs. Liquid-cooled systems also allow higher GPU density per rack, reducing interconnect latency and physical footprint. Every major new AI rack (DGX GB200, MI350X cluster) ships liquid-cooled by default.
Known deployments
19 disclosed deployments · who is building what
| Operator | System |
|---|---|
OC Oracle Cloud | |
All systems
10 systems · 6 manufacturers
Google's high-performance TPU pod for large-scale training. 8,960 TPU v5p chips in a single pod connected via ICI 3.0 fabric. Powers Gemini ...
Next-gen liquid-cooled rack with Blackwell Ultra GPUs. 72x B300 GPUs with 288 GB HBM3e per GPU (vs 192 GB on B200). Designed for reasoning-h...
NVIDIA's flagship liquid-cooled rack. 72 Blackwell GPUs + 36 Grace CPUs connected via NVLink 5.0 in a single 72-GPU domain. Designed for tri...
Google's latest custom AI accelerator in pod configuration. 256 TPU v6e chips connected via custom ICI (Inter-Chip Interconnect). Optimized ...
Microsoft's first custom AI silicon at rack scale. Maia 100 chips fabricated at TSMC on N5 with HBM3e. Designed for Azure AI inference and f...
8-GPU Blackwell node for enterprises that don't need the full NVL72 rack. Air-cooled with NVLink 4.0 interconnect. The workhorse for inferen...
AMD's latest 8-GPU OAM platform with MI325X accelerators. 256 GB HBM3e per GPU for the largest memory footprint in its class. Infinity Fabri...
AWS custom silicon training server. 16x Trainium2 chips in a single UltraServer node connected via NeuronLink. Designed to compete with NVID...
The system that launched the AI infrastructure boom. 8x H100 SXM GPUs connected via NVLink 4.0. Still the most widely deployed AI training s...
Wafer-scale AI compute appliance. A single CS-3 contains one WSE-3 chip (the largest chip ever made, using an entire 300mm wafer). 44 GB of ...
Frequently asked
Computed from the dataset on this page · see its as-of date
What is a DGX GB200 NVL72?
The DGX GB200 NVL72 is NVIDIA's rack-scale AI system: 72 Blackwell GPUs and 36 Grace CPUs in one liquid-cooled rack, linked by NVLink. The dataset lists 720 PFLOPS of FP8 compute and 120 kW of rack power. Known deployments include Microsoft Azure, Oracle Cloud, CoreWeave and xAI.
What does TCO per PFLOPS mean?
Total cost of ownership per PFLOPS per year puts different AI systems on one scale. BenchGecko adds the hardware cost amortized over 3 years to the power cost at $0.05/kWh, then divides by the system's FP8 compute in PFLOPS. Lower is better.
Why do some systems require liquid cooling?
Current AI accelerators draw up to 1,200 W each in the dataset, and the DGX GB300 NVL72 rack draws 140 kW. Air cooling cannot remove that much heat efficiently. Liquid cooling (direct-to-chip or immersion) allows denser racks and a lower PUE: 6 of the 10 tracked systems are liquid cooled.
How does a TPU pod compare to an NVIDIA DGX rack?
They are different architectures. The TPU v5p Pod links 8,960 TPU chips in one domain over Google's ICI fabric, while NVIDIA's NVL72 is a 72-GPU rack. TPU pods are only available on Google Cloud; DGX systems are sold by many OEMs for on-premise deployment.
What is the Cerebras CS-3 and why is it different?
The Cerebras CS-3 is built around a single wafer-scale chip, the WSE-3, which uses a whole silicon wafer. Instead of HBM stacks it has 44 GB of on-chip SRAM, which removes the memory bandwidth bottleneck of conventional GPU systems. The tradeoff is a higher per-system cost and a different programming model.
Why are hyperscalers building custom silicon?
Google (TPU), AWS (Trainium), Microsoft (Maia) and Meta (MTIA) build their own AI chips to reduce dependence on NVIDIA, tune for their own workloads and lower cost at scale. Custom silicon also diversifies supply. Systems data as of Apr 14, 2026.
See also
Keep exploring the compute graph