What it is
Every time you send a prompt to a model, the provider has to read all of it before it writes a single word of reply. If most of your prompts begin the same way, with the same instructions, the same documents or the same tool definitions, that is a lot of repeated reading. Prompt caching lets the provider keep the result of that reading and reuse it.
Think of a barista who remembers your usual order. They still make a fresh coffee every time, but they skip the part where you explain what you want.
How it works
When a model reads a prompt, it builds an internal record of it. Prompt caching saves that record for a short time. When your next request starts with exactly the same text, the model skips the saved part and only processes what is new. The reply is still generated fresh, so you can get a different answer each time. Only the reading is reused, never the output.
Two details matter:
- The match must be exact, and it starts at the beginning. A single changed word near the top breaks the match for everything after it. Put stable content first (instructions, reference documents, examples) and the part that changes last (the user's question).
- Caches expire. Providers keep them for minutes, not days, unless you pay for a longer lifetime. A cache you do not touch goes away.
Providers differ. Some apply caching automatically above a minimum prompt length. Others make you mark where the cacheable part ends. Some charge a little extra the first time the prompt is stored, then a steep discount on every token read from the cache afterward. Check your provider's current documentation for exact thresholds, lifetimes and prices, because they change.
Why it matters to you
If you send the same long material over and over, caching is often the biggest single cut to your bill. Typical cases:
- A support bot that carries a long policy document in every request.
- A coding assistant that re-reads the same files on each turn.
- An agent that loops through many steps, each one repeating the full history and its list of tools.
Caching also lowers the wait before the first word appears, because there is less to read. It makes long prompts far more practical to use, since a large context window is cheaper to fill when most of it is cached.
What it is not
Prompt caching is not the same as saving answers. A "semantic cache" returns a stored reply when a new question looks similar to an old one. Prompt caching never does that. It only speeds up the reading step of inference.
A quick test
Look at the usage numbers your API returns. Most providers report how many input tokens were read from the cache. If that number is zero on your second request, something at the top of your prompt is changing, such as a timestamp or a reordered list.