GPU
Supply-constrained and margin-rich.
A GPU is the accelerator that trains and serves most large AI models.
“A GPU is the accelerator that trains and serves most large AI models.”
Read full chapterSix terms for the angry question at the VC meeting.
Supply-constrained and margin-rich.
A GPU is the accelerator that trains and serves most large AI models.
“A GPU is the accelerator that trains and serves most large AI models.”
Read full chapterThe memory bottleneck.
HBM3e (High Bandwidth Memory 3e (Enhanced)) is a JEDEC HBM memory generation from 2024, with 1,180 GB/s of bandwidth per stack.
How compute capacity gets measured.
FP16 is a 16-bit floating point format widely used in neural network training and inference.
“FP16 is a 16-bit floating point format widely used in neural network training and inference.”
Read full chapterWhy frontier training runs are so costly.
The phase where a model learns from internet-scale text via next-token prediction · typically 15-30T tokens and millions of GPU hours.
“Pretraining is the moat.”
Read full chapterWhy serving costs never go to zero.
The tokens the model writes back to you · priced 3-5× higher than input because decoding is sequential.
“Output token economics drove every major model architecture decision since 2024. MoE, reasoning tiers, speculative decoding · all optimizing one thing.”
Read full chapterThe spending that all this requires.
The billions hyperscalers and AI labs spend each year on GPUs, datacenters, and training clusters · the #1 driver of AI spending narrative.
“Nobody knows yet · but BenchGecko tracks it daily.”
Read full chapterBy the end you can explain AI capex, HBM bottlenecks, training and inference economics · and how to read the AI Bubble Index.
Seven terms that decode whether AI is overpriced, fairly priced, or criminally underpriced. Read in order.