Input Tokens
The tokens in your prompt · billed per million, typically 3-5× cheaper than output tokens.
Text reviewed October 5, 2026
The tokens in your prompt · billed per million, typically 3-5× cheaper than output tokens.
Basic
When you call an LLM API, everything you send (system prompt + user message + retrieved docs + conversation history) counts as input tokens. Providers charge per million input tokens. Current prices for every tracked model are on the BenchGecko pricing pages.
Deep
Input token cost scales with prompt length. Long system prompts + full retrieval + long conversation history = expensive. Most production apps land at 70-85% of total token cost on input despite paying more per output token, because inputs are typically 5-20× longer than outputs. Current prices for every tracked model are on the BenchGecko pricing pages.
Expert
Input tokens are processed in parallel during the prefill phase · compute-bound, not memory-bound. This is why input is cheap per token relative to output · you can batch prefill work efficiently across many requests. Cost accounting: the cheapest way to run long-context workloads is aggressive prompt caching + retrieval-pruned context. Cache-aware architectures (stable system prompt + volatile user context) let you pay 10× less for the same effective context.
Depending on why you're here
- ·The stuff you send to the AI · shorter is cheaper
- ·Usually cheaper than what the AI sends back
- ·Why long conversations get expensive over time
- ·Keep prompts tight · long context bills you per call
- ·Cost most wars are won or lost on input token discipline
- ·Input margin is structurally thin · providers recoup on output
- ·Prompt caching can reshape competitive pricing positioning
- ·Processed in parallel during prefill · compute-bound
- ·5-20× longer than outputs in typical workloads
Input token discipline separates teams that can scale from teams that can't. Cache aggressively.