Post-training is everything done to an AI model after its main training run is finished. Pretraining teaches a model to predict the next piece of text across a huge pile of data. Post-training shapes that raw ability into something that follows instructions, behaves safely or does one job well.
The usual steps
- Supervised fine-tuning. The model is trained on curated examples of good answers, such as questions paired with ideal replies, so it learns the format and style you want.
- Learning from feedback. People or other models rank the model's answers, and the model is adjusted toward the preferred ones. This is how many assistants learn to be helpful and to decline harmful requests.
- Reinforcement learning on checkable tasks. For problems with a verifiable answer, like math or code that passes tests, the model is rewarded when it gets them right. This is a major way reasoning models learn to work through a problem in a long chain of thought.
- Task specialization. A general model is trained further to do one narrow job, such as scoring a fixed set of choices or writing in a particular domain.
- Safety training. The model is trained to refuse certain requests and to follow its developer's rules, which is part of AI alignment work.
An analogy
Pretraining is a general education. Post-training is job training. The graduate knows a lot about everything, and the training teaches them how to behave at work, what the boss expects and how to do one role well.
Why it matters to you
Two models built on the same pretrained base can feel completely different because of post-training. It explains why a small, specialized model can beat a larger general one on a narrow task, and why a model's refusals, tone and reliability come mostly from this stage, not from its size.
It also explains a naming pattern you will see: many products are an existing open-weight model post-trained for a specific purpose. If you rely on one, ask what the base model was and what testing the post-training received, since safety work done on the base may not carry over.
How it relates to other terms
Post-training is not inference, which is running a finished model to get answers. It is also broader than fine-tuning, which is one technique inside it. Distillation, where a small model learns from a larger one, is another method that often appears at this stage.