PricingReading · ~3 min · 57 words deep

Input Tokens

The tokens in your prompt · billed per million, typically 3-5× cheaper than output tokens.

Text reviewed October 5, 2026

TL;DR

The tokens in your prompt · billed per million, typically 3-5× cheaper than output tokens.

Level 1

When you call an LLM API, everything you send (system prompt + user message + retrieved docs + conversation history) counts as input tokens. Providers charge per million input tokens. Current prices for every tracked model are on the BenchGecko pricing pages.

Level 2

Input token cost scales with prompt length. Long system prompts + full retrieval + long conversation history = expensive. Most production apps land at 70-85% of total token cost on input despite paying more per output token, because inputs are typically 5-20× longer than outputs. Current prices for every tracked model are on the BenchGecko pricing pages.

Level 3

Input tokens are processed in parallel during the prefill phase · compute-bound, not memory-bound. This is why input is cheap per token relative to output · you can batch prefill work efficiently across many requests. Cost accounting: the cheapest way to run long-context workloads is aggressive prompt caching + retrieval-pruned context. Cache-aware architectures (stable system prompt + volatile user context) let you pay 10× less for the same effective context.

The takeaway for you
If you are a
Curious · Normie
  • ·The stuff you send to the AI · shorter is cheaper
  • ·Usually cheaper than what the AI sends back
  • ·Why long conversations get expensive over time
If you are a
Builder
  • ·Keep prompts tight · long context bills you per call
  • ·Cost most wars are won or lost on input token discipline
If you are a
Investor
  • ·Input margin is structurally thin · providers recoup on output
  • ·Prompt caching can reshape competitive pricing positioning
If you are a
Researcher
  • ·Processed in parallel during prefill · compute-bound
  • ·5-20× longer than outputs in typical workloads
Gecko's take

Input token discipline separates teams that can scale from teams that can't. Cache aggressively.

Prompt caching (Claude, OpenAI), retrieval-pruned RAG context, shorter system prompts, tighter few-shot examples.