What it is
Building an AI model has two big phases. Training is the long, expensive process where the model learns from data. Inference is everything after: using the finished model to get an answer. Every time you send a prompt to a chatbot and it replies, that is inference.
An analogy: training is a chef spending years learning to cook. Inference is the chef making your dinner tonight.
What happens during inference
Your prompt is split into tokens and fed through the model, which predicts the next token, then the next, until the answer is complete. The model's weights do not change while this happens. It is applying what it already learned, not learning from your chat.
Why it matters to you
Inference is where most of the practical questions live.
- Cost. Providers charge for the tokens processed. Your bill is an inference bill.
- Speed. Two numbers matter: how long until the first token appears, and how fast tokens stream after that. Bigger models and longer prompts are generally slower.
- Capacity. Serving millions of people at once needs a huge amount of specialized hardware, which is why providers set rate limits and usage caps.
- Limits. The context window caps how much the model can consider in one inference call.
Ways to make it cheaper or faster
Engineers use several tricks: smaller models for easy tasks, caching repeated parts of a prompt, batching many requests together, and compressing models. Architectures like mixture of experts lower the compute per token by using only part of the network at a time.
Reasoning and inference cost
Some models spend extra effort thinking before they answer, producing hidden intermediate tokens. That improves results on hard problems but raises inference cost and delay, so it is worth using only where the task needs it.
Where it runs
Inference can run on a provider's servers through an API, or on your own hardware if you have an open-weight model. Running it locally trades convenience for control.
The takeaway
When you compare AI options, compare inference: cost per task, speed, and limits. Training happens once. Inference happens every time you press enter.