Skip to content
LiveNext 11:37:01

Definition

What is Inference?

Inference is Inference is the act of running a trained AI model to produce an output, such as answering a prompt, as opposed to training the model in the first place.

What it is

Building an AI model has two big phases. Training is the long, expensive process where the model learns from data. Inference is everything after: using the finished model to get an answer. Every time you send a prompt to a chatbot and it replies, that is inference.

An analogy: training is a chef spending years learning to cook. Inference is the chef making your dinner tonight.

What happens during inference

Your prompt is split into tokens and fed through the model, which predicts the next token, then the next, until the answer is complete. The model's weights do not change while this happens. It is applying what it already learned, not learning from your chat.

Why it matters to you

Inference is where most of the practical questions live.

  • Cost. Providers charge for the tokens processed. Your bill is an inference bill.
  • Speed. Two numbers matter: how long until the first token appears, and how fast tokens stream after that. Bigger models and longer prompts are generally slower.
  • Capacity. Serving millions of people at once needs a huge amount of specialized hardware, which is why providers set rate limits and usage caps.
  • Limits. The context window caps how much the model can consider in one inference call.

Ways to make it cheaper or faster

Engineers use several tricks: smaller models for easy tasks, caching repeated parts of a prompt, batching many requests together, and compressing models. Architectures like mixture of experts lower the compute per token by using only part of the network at a time.

Reasoning and inference cost

Some models spend extra effort thinking before they answer, producing hidden intermediate tokens. That improves results on hard problems but raises inference cost and delay, so it is worth using only where the task needs it.

Where it runs

Inference can run on a provider's servers through an API, or on your own hardware if you have an open-weight model. Running it locally trades convenience for control.

The takeaway

When you compare AI options, compare inference: cost per task, speed, and limits. Training happens once. Inference happens every time you press enter.

Questions people ask

What is the difference between training and inference?

Training teaches the model from data and changes its weights. Inference uses the finished model to produce outputs and does not change the weights.

Does the model learn from my chat during inference?

No. The weights stay fixed. Anything it appears to remember comes from the conversation held in its context window.

Inference in the news