This conversation is happening in many finance meetings. The engineering lead says the AI model is cheaper than last year, and it is. The finance lead says the invoice is higher, and it is. Both are right, and the gap is worth understanding.
The price really is falling
Epoch AI found that the price of getting GPT-4-level results on a set of PhD-level science questions fell by about 40 times a year. Across the tasks it measured, the decline ranged from 9 to 900 times a year, though Epoch cautions that the fastest drops are recent and may not last. In March 2026, Gartner predicted that by 2030, running a trillion-parameter model will cost AI providers more than 90% less than in 2025. So why are bills going the other way?
Your bill has three parts, and only one is falling

Your AI bill is the price per token, times the tokens each task uses, times the number of tasks you run. (A token is a small piece of text, about four characters.) The first number is falling fast. The other two are rising faster.
Tokens per task are going up. A chatbot reads a question and answers. An AI agent plans, calls tools, reads the results, checks its work and tries again. Gartner's March release noted that agentic models need 5 to 30 times more tokens per task than a standard chatbot, and its August 2026 forecast expects AI inference costs per agentic workflow to rise more than fivefold through 2028. Gartner calls this the inference paradox: "The rate of innovation is outpacing the cost curve."
The number of tasks is going up too. The pilot that answered internal questions now drafts proposals, summarizes calls and checks contracts. As Gartner analyst Will Sommer put it, "Product leaders cannot rely on more efficient token economics to rationalize AI costs."
Round numbers, for illustration: a chatbot answers a question about a late order in one model call. An agent plans, looks up the order, checks the policy, drafts a reply and reviews it, in five calls that each re-read the conversation: easily fifteen times the tokens, for a much better answer. Halve the token price and each request still costs seven and a half times more. Route three times as many requests to it, and that feature's bill is over twenty times what it was. Every decision made sense, which is why the bill surprises people.
Where the tokens actually go
- Context resent on every turn. Instructions, history and documents go with every request.
- Agents that loop. Ten retries of a failing step means paying for ten attempts.
- Hidden reasoning. A reasoning model's thinking is usually billed as output, even unseen.
- Retrieval that grabs too much. Search passes far more text than the answer needs.
- Tool results pasted back in full. A whole web page or log file goes back when one line was needed.
- One model for everything. Sending simple classification to the most powerful model is like sending every letter by overnight courier.
Seven ways to bring the bill under control
In the FinOps Foundation's State of FinOps 2026 report, 98% of FinOps teams said they now manage AI spending, up from 63% in 2025 and 31% in 2024, and AI cost management is the skill they most want to add.
- Measure cost per outcome, not per token. The cost of each resolved ticket or finished task shows whether AI pays for itself.
- Route work to the right model. Use small models for routine, high-volume tasks. Gartner advises that frontier-level inference "must be heavily gated" and saved for complex reasoning.
- Trim the context. Retrieve fewer, better passages and summarize long histories.
- Cache what repeats. Many providers charge less for reused, cached prompt content.
- Put limits on agents. Cap steps, retries and tokens per task.
- Batch work that isn't urgent. Overnight and bulk jobs cost less in batch.
- Give every AI feature a budget and an owner. Set alerts and assign costs to the teams that create them.
The real question
Cheaper tokens don't manage your costs for you. Ask your team or vendor what one completed task costs now versus three months ago, which model handles each step and why, what stops an agent from looping, and who sees the AI bill each month. The organizations that get this right design for efficiency, measure what each outcome costs, and use the least expensive model that does the job well.
If your AI costs are climbing faster than their value, our data and AI team can help you find where the tokens go and spend them wisely.
