Tokens
The fundamental unit that LLMs read and generate · 1 token ≈ 0.75 English words or 4 characters.
Text reviewed October 5, 2026
The fundamental unit that LLMs read and generate · 1 token ≈ 0.75 English words or 4 characters.
Basic
Every LLM breaks text into tokens before processing. A token is usually a subword fragment, not a whole word · "unbelievable" might be ["un", "believ", "able"]. Pricing, context windows, and speed metrics are all measured in tokens. 1M tokens ≈ 750K English words ≈ a mid-length novel.
Deep
Tokenization is language-specific. English averages 1.3 tokens per word. Code averages 2.5. Chinese and other non-Latin scripts average 2-4× more tokens per character than English, which is why non-English API calls cost more. Tokenizer algorithms vary: GPT uses BPE (Byte Pair Encoding), Llama uses SentencePiece, Claude uses a Claude-specific tokenizer. Larger vocab means fewer tokens per input but larger embedding tables. See the related terms and live BenchGecko data for current examples.
Expert
BPE splits text greedily by merging the most frequent byte pairs until the vocabulary is full. SentencePiece works at the raw byte level and is language-agnostic. Tiktoken (OpenAI) exposes a fast Rust-backed encoder. Token count = cost. Multilingual efficiency is a known challenge: Japanese costs ~4× more per character than English on most providers.
Depending on why you're here
- ·AI reads and writes in tokens, not words
- ·1 token is about 3/4 of a word
- ·Why the bill jumps when you send long messages
- ·Count tokens before sending · use tiktoken or the provider SDK
- ·Budget for 1.3 tokens per English word, 2.5 per code word
- ·Non-English workloads are more expensive per character
- ·Token pricing is the unit economic · every benchmark comparison normalizes to $/M tokens
- ·Tokenizer efficiency is a hidden lever in multilingual cost
- ·Model providers occasionally revise tokenizers · watch for pricing shifts
- ·BPE, SentencePiece, Tiktoken are the common algorithms
- ·Vocabulary size trades off embedding table size vs tokens per input
- ·Non-Latin scripts suffer 2-4× worse tokenization ratios
Tokens are the units of AI billing. Understand them or overpay.