MMLU
The baseline knowledge benchmark everyone cites.
MMLU is a knowledge benchmark tracked on BenchGecko. Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.
Six terms to go from "I need an AI" to "here is the cheapest model that meets my spec."
The baseline knowledge benchmark everyone cites.
MMLU is a knowledge benchmark tracked on BenchGecko. Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.
If your workload is code, this is the one to care about.
A benchmark where models attempt real GitHub issues · judged by whether their patch passes the project's test suite.
“SWE-bench is the single most-watched AI benchmark of 2026. Every coding agent release ships a SWE-bench number first.”
Read full chapterHow much input the model can hold at once.
The max number of tokens · input + output · a model can handle in a single request.
“Context windows hit diminishing returns past 200K for most workloads. 1M+ is for agents and codebase-scale retrieval, not chat.”
Read full chapterPrice per million input tokens, the biggest line on most bills.
The tokens in your prompt · billed per million, typically 3-5× cheaper than output tokens.
“Input token discipline separates teams that can scale from teams that can't. Cache aggressively.”
Read full chapterSpeed: tokens per second, which shapes the user experience.
The total tokens-per-second a serving cluster handles across ALL concurrent requests · vs tokens-per-second of a single request.
“Throughput per dollar is the number that actually matters. Every pricing war is a throughput war in disguise.”
Read full chapterHow a model connects to your code and tools in production.
The API mechanism that lets models invoke external tools by emitting structured JSON that matches a schema.
“Function calling is solved infrastructure. The interesting problems moved up the stack · to orchestration, memory, planning.”
Read full chapterBy the end you can evaluate a model by benchmark match, price, context window, and speed · and pick the winner for your specific workload.
Seven terms that decode whether AI is overpriced, fairly priced, or criminally underpriced. Read in order.