← home

Token prices fell. The AI budget still went up.

August 27, 2026

By the LLM CFO team

Lower token prices are not the same as lower AI spend. A budget can rise after a major provider discount because usage, model mix, context length, and agent activity grew faster than the rate fell. The finance question is not whether AI became cheaper per token; it is what changed between the approved plan and the invoice.

Build the variance bridge

Start with the locked baseline: provider invoices, approved traffic assumptions, and the model mix used in the budget. Explain the movement in four steps. Rate captures price changes and negotiated discounts. Volume captures requests and tokens. Mix captures movement between models, providers, regions, and workloads. Efficiency captures the cost of delivering one successful outcome, including retries, long context, tool calls, and unused output.

VarianceFinance interpretationOwner
RateSupplier economics changed.Procurement
VolumeDemand exceeded the plan.Product and sales
MixThe business used a different capability.Engineering
EfficiencyThe same outcome now costs more.Engineering and product

Do not reward the wrong number

A falling blended price can conceal a deteriorating unit margin. A product may be acquiring customers faster while each completed workflow requires more reasoning, retrieval, or retries. Conversely, a higher bill can be healthy if it bought revenue-producing usage. The board metric should pair spend with an outcome: cost per successful task, gross margin per AI feature, or AI cost per active customer segment.

What the monthly close needs

Finance should receive a provider-reconciled spend statement, a variance bridge, and a forecast using actual usage drivers. Ask for p50 and p95 cost per successful task, not just a company-wide average. Require owners for the three largest movements and record whether each is temporary demand, a deliberate investment, or an avoidable leak.

The budget decision

Do not respond to a variance by applying a universal token cap. Set guardrails by workload: latency-sensitive customer requests, background jobs, evaluation traffic, and autonomous agents have different economic limits. Approve more spend when the contribution margin supports it; require a routing, caching, or stopping-rule change when efficiency deteriorates without a corresponding business benefit.

AI price cuts create room for better economics, but only if the organization can see where the room went. A disciplined close turns a surprising invoice into a decision about demand, capability, and value.

Related

Related

← Back to llmcfo.com