Find out what this week's small-model price cuts do to your bill
› I run [describe your app, agent or workflow] on [current model and its price per million input and output tokens]. A typical month looks like this: [number of requests], an average prompt of [tokens], an average reply of [tokens], and [share]% of prompts go over 100,000 tokens.
Two small models now list the same base price of $0.10 per million input tokens and $0.50 per million output tokens. Option A rises to $0.50 input and $2.50 output once a prompt passes 100,000 tokens. Option B rises to $0.20 input and $0.75 output once a prompt passes 272,000 tokens. Assume Option A may use up to 25% more tokens for the same text because of a newer tokenizer.
1. Calculate my monthly cost on my current model, Option A and Option B. Show the math in a table.
2. Tell me the prompt length at which the cheaper option flips.
3. Suggest three changes to my prompts, caching or routing that would cut the bill further, with the rough saving for each.
Before you start, ask me for any number above I left blank. Do not guess my usage.
Works in any AI chat. In Magai, run it against several models at once.