What a token is
Language models do not read letters or whole sentences. Before your text reaches the model, a tokenizer cuts it into pieces called tokens. A token can be a whole common word like "the", part of a longer word like "un" plus "believ" plus "able", a single punctuation mark, or a few characters of code.
Think of tokens as the Lego bricks of text. The model never sees your sentence as one object. It sees a row of bricks, each with an ID number, and it predicts the next brick, then the next.
How big is a token?
It depends on the tokenizer and the language, so any rule of thumb is rough. In ordinary English, a token is often around three quarters of a word, which means 100 words come to roughly 130 tokens. Code, numbers, and languages other than English usually split into more tokens for the same meaning. If you need an exact count, use the tokenizer or token counter your model provider publishes, because different model families count differently.
Why tokens matter to you
Tokens are the unit almost everything else is measured in.
- Price. AI providers usually charge per million tokens, with one rate for what you send (input) and often a different rate for what the model writes back (output).
- Limits. A model's context window is counted in tokens. Your prompt, any documents you paste in, the conversation so far, and the model's reply all share that budget.
- Speed. Models produce output one token at a time, so a longer answer takes longer.
A practical example: pasting a 50-page report into a chat spends a large share of your token budget before the model writes a word. Trimming what you send is the easiest way to cut cost and keep the model focused.
Common confusions
A token is not a word, and it is not a character. It also has nothing to do with the login tokens used for security, or with cryptocurrency. Same word, different idea.
Tokens are the currency of inference, the work of running a trained model to get an answer. Models built as a mixture of experts still count tokens the same way, even though only part of the model works on each one.
The takeaway
When someone quotes an AI price, a limit, or a speed, ask what it is per token and which kind of token, input or output. That one question explains most AI bills.