Token prices fell. The AI budget still went up.
August 27, 2026
Lower token prices are not the same as lower AI spend. A budget can rise after a major provider discount because usage, model mix, context length, and agent activity grew faster than the rate fell. The finance question is not whether AI became cheaper per token; it is what changed between the approved plan and the invoice.
Build the variance bridge
Start with the locked baseline: provider invoices, approved traffic assumptions, and the model mix used in the budget. Explain the movement in four steps. Rate captures price changes and negotiated discounts. Volume captures requests and tokens. Mix captures movement between models, providers, regions, and workloads. Efficiency captures the cost of delivering one successful outcome, including retries, long context, tool calls, and unused output.
| Variance | Finance interpretation | Owner |
|---|---|---|
| Rate | Supplier economics changed. | Procurement |
| Volume | Demand exceeded the plan. | Product and sales |
| Mix | The business used a different capability. | Engineering |
| Efficiency | The same outcome now costs more. | Engineering and product |
Do not reward the wrong number
A falling blended price can conceal a deteriorating unit margin. A product may be acquiring customers faster while each completed workflow requires more reasoning, retrieval, or retries. Conversely, a higher bill can be healthy if it bought revenue-producing usage. The board metric should pair spend with an outcome: cost per successful task, gross margin per AI feature, or AI cost per active customer segment.
What the monthly close needs
Finance should receive a provider-reconciled spend statement, a variance bridge, and a forecast using actual usage drivers. Ask for p50 and p95 cost per successful task, not just a company-wide average. Require owners for the three largest movements and record whether each is temporary demand, a deliberate investment, or an avoidable leak.
The budget decision
Do not respond to a variance by applying a universal token cap. Set guardrails by workload: latency-sensitive customer requests, background jobs, evaluation traffic, and autonomous agents have different economic limits. Approve more spend when the contribution margin supports it; require a routing, caching, or stopping-rule change when efficiency deteriorates without a corresponding business benefit.
AI price cuts create room for better economics, but only if the organization can see where the room went. A disciplined close turns a surprising invoice into a decision about demand, capability, and value.