Learning path6 terms · ~18 min read

Why AI is Expensive

Six terms for the angry question at the VC meeting.

Start · GPU
ChipsChapter 1 of 6

Supply-constrained and margin-rich.

TL;DR

A GPU is the accelerator that trains and serves most large AI models.

“A GPU is the accelerator that trains and serves most large AI models.”

Read full chapter
MemoryChapter 2 of 6

The memory bottleneck.

TL;DR

HBM3e (High Bandwidth Memory 3e (Enhanced)) is a JEDEC HBM memory generation from 2024, with 1,180 GB/s of bandwidth per stack.

Read full chapter
ConceptsChapter 3 of 6

How compute capacity gets measured.

TL;DR

FP16 is a 16-bit floating point format widely used in neural network training and inference.

“FP16 is a 16-bit floating point format widely used in neural network training and inference.”

Read full chapter
ConceptsChapter 4 of 6

Why frontier training runs are so costly.

TL;DR

The phase where a model learns from internet-scale text via next-token prediction · typically 15-30T tokens and millions of GPU hours.

“Pretraining is the moat.”

Read full chapter
PricingChapter 5 of 6

Why serving costs never go to zero.

TL;DR

The tokens the model writes back to you · priced 3-5× higher than input because decoding is sequential.

“Output token economics drove every major model architecture decision since 2024. MoE, reasoning tiers, speculative decoding · all optimizing one thing.”

Read full chapter
EconomyChapter 6 of 6

The spending that all this requires.

TL;DR

The billions hyperscalers and AI labs spend each year on GPUs, datacenters, and training clusters · the #1 driver of AI spending narrative.

“Nobody knows yet · but BenchGecko tracks it daily.”

Read full chapter
What you learned

By the end you can explain AI capex, HBM bottlenecks, training and inference economics · and how to read the AI Bubble Index.

Keep learning
Next path · 7 terms
The AI Bubble Explained

Seven terms that decode whether AI is overpriced, fairly priced, or criminally underpriced. Read in order.