What it is
A tokenizer is the piece of software that turns your text into tokens before a model reads it, and turns the model's tokens back into text afterward. Models do not see letters or words. They see a sequence of numbers, each standing for one token, and the tokenizer is the dictionary that maps between the two.
How it works
Most modern tokenizers learn their vocabulary from large amounts of text. Common words get a single token. Rarer words are split into pieces. A word like "unbelievable" might become "un", "believ" and "able", while a very common word like "the" is one token on its own.
The vocabulary is fixed when the model is built. Each model family usually has its own tokenizer, so the same sentence can produce a different number of tokens on different models.
Why it matters to you
The tokenizer quietly sets your bill and your limits, because AI usage is priced and capped in tokens. A few practical effects:
- Cost: if one tokenizer turns your text into 20 percent more tokens than another, the same prompt costs 20 percent more at an identical per-token price. When comparing providers, compare cost on your own text, not just the price per million tokens.
- Context limits: your context window is counted in tokens, so a less efficient tokenizer fits less text into the same window.
- Languages: tokenizers trained mostly on English often split other languages, and also code, numbers and emoji, into more tokens. The same message in another language can cost noticeably more.
- Odd behavior: because a model sees tokens rather than letters, tasks like counting the letters in a word or reversing a string can trip it up.
An example
Take the sentence "Tokenizers split text." One tokenizer might produce five tokens and another seven. If you send ten million such sentences a month, that gap shows up directly on your invoice, even though the text never changed.
How to check
Providers publish token counters or return token counts with each response. Paste a representative sample of your real prompts into the counter for each model you are weighing. A rough rule of thumb for English is that a token is about three quarters of a word, but your own samples are better than any rule.
When a provider updates its model, it sometimes updates the tokenizer too, which can change counts for the same input. That is worth checking before you compare old and new costs. Reusing a long, repeated prompt prefix through prompt caching lowers the cost of those tokens, and the work of producing a reply is called inference.