My electric bill last month was about $1,000. PG&E can tell me exactly how many kilowatt-hours I consumed, when I consumed them, and what each one cost. What the bill can’t tell me is whether that money went to the air conditioner, the kids’ video games, or a neighbor with an extension cord quietly mining crypto off my outlet. Same meter, same accuracy, three very different conversations about what to do next.
Most enterprise AI cost reporting works exactly like that bill. The dashboards are precise about consumption: tokens per user, spend per model, cost per department. They say nothing about purpose. The FinOps Foundation’s 2026 State of FinOps report found 73 percent of organizations exceeded their AI budget projections. And the visibility executives believe they have doesn’t survive a simple test. When WitnessAI surveyed 300 senior decision-makers this spring, 96 percent said they were at least mostly confident in their visibility into the AI tools, models, and agents touching company data. Only 14 percent could produce a complete inventory within a day, and a third would need more than a week. That’s the PG&E kind of visibility: confidence in the meter, with no idea which of the three conversations you should be having.
I think the fix is a discipline I’d call behavioral FinOps: attaching intent to every token, whether a person or an agent spent it, on both the observability side and the control side. Seeing what each dollar of inference was for, and acting on that understanding in the traffic path, before the token is spent.
What spend looks like when you can see intent
Token metering answers how much. Behavior answers what for, and every dollar of recoverable waste sits behind that second question.
When every prompt is classified by intent before it reaches a model, the meter reading turns into an itemized bill with a purpose column. Business use separates from personal use, by user and by department. Enterprise seats running at a fraction of their token allotment show up next to users on cheaper tiers who hit their ceilings every week. And the workloads behind the spend become visible: customer support, code generation, contract review, vacation planning.
That last one matters more than it sounds. Industry analysis puts up to 70 percent of generic AI chatbot traffic in the personal or non-work category (Index.dev, 2025). On enterprise licenses billed per token, the company absorbs all of it.
The deeper value of intent data is that it tells you which problem you actually have. Non-productive spend is not one category, and my electric bill turns out to be a decent taxonomy for it. The air conditioner is legitimate work running on an expensive resource: a marketing analyst reformatting text through a frontier model when a lightweight one produces identical output. The kids’ video games are personal use of company resources: harmless prompts, but the tokens aren’t free. The neighbor’s extension cord is deliberate misuse: an outside engineer redirecting your public customer-service chatbot into solving coding challenges, burning your inference budget for their benefit.
A spend dashboard without behavior renders all three as the same line item. It tells you the bill went up. It can’t tell you whether you have a routing problem, a policy problem, or an abuse problem, and those have three different fixes.
The outcome side of the ledger is just as blind. In the same survey, only 9 percent of executives said more than three quarters of their AI initiatives have delivered measurable financial return. The board is asking what the spend bought. A meter can’t answer that either.
And a growing share of the spend has no human behind it at all. An agent is the smart-home controller with my credit card on file: it can switch on any appliance itself, and it runs all night at machine speed. Agents chain dozens of model calls per task and resend their accumulated context at every step. That’s why Gartner puts agentic workloads at 5 to 30 times the token cost per task of a standard chatbot. The same firm’s August 2026 press release predicted inference costs per agentic workflow will rise more than fivefold through 2028. A widely reported 2026 incident saw a LangChain multi-agent system run an infinite loop for 11 days and burn $47,000 in API charges. In many enterprises the costliest AI users are no longer people. A busy agent and a productive agent look identical on the bill, and the difference between them is the purpose behind each call, which a meter doesn’t record.
Control that acts on behavior
Observability is only half the discipline. If PG&E itemized my bill by behavior, the fixes would be obvious, and I’d make them the same afternoon. Disconnect the extension cord. Put hours on the video games. Adjust the thermostat schedule. Knowing which behavior each dollar is tied to is what turns a bill into a set of decisions. AI spend works the same way, except the enforcement can run in the traffic path automatically, where each of the three patterns gets its own response.
Deliberate misuse gets cut off. When intent classification recognizes that a prompt to your customer-facing chatbot is a coding challenge rather than a customer question, it blocks the request before it reaches paid inference and alerts security. B2C chatbot abuse is a cost problem and a security problem at once, and it never shows up on a per-user spend report because the user isn’t yours.
Personal use gets a house rule. It triggers a configurable response: warn the employee, redirect them to a personal tool, or block, depending on your policy. Blocking everything kills adoption, which is its own cost. Classifying everything lets you set the policy deliberately instead of absorbing the spend by default.
Misrouted work gets a cheaper model. Prompts are scored for risk, complexity, and purpose, then routed to the cheapest model that clears the quality bar, and that applies to agents as much as to people. An agent that takes twenty steps to finish a task multiplies whatever routing policy you set twenty times, which makes agent traffic the highest-return place to route on behavior. The published price gap between a frontier model and a lightweight one, GPT-4o versus GPT-4o-mini for example, is 94 percent. The RouteLLM research presented at ICLR 2025 showed intelligent routing cuts total inference cost 40 to 80 percent with no measurable quality loss on routine work. Sensitive prompts route to secure internal models for the same reason, because routing on behavior serves risk and cost with one mechanism.
None of these controls is possible from a billing console. By the time the token shows up on an invoice, every one of those decisions has already been made, badly, by default.
Where the market falls short
A category of products has emerged to meter AI spend: gateway layers and cost dashboards that sit between you and your model providers and count tokens with real precision. Counting is necessary. Consolidated metering across providers is useful, and most enterprises lack even that.
But a meter without behavior has a low ceiling. On the observability side, it produces a precise accounting of consumption, itemized by user and model but silent on purpose. On the control side, a tool that can’t classify intent can only allow / block / route by user or agent. It can’t block the outsider abusing your chatbot with irrelevant prompts. It can’t route the reformatting job to a cheaper model because it can’t tell reformatting from contract analysis. The savings that routing and filtering produce are structurally out of reach.
The practical test is simple, and it’s the first thing I’d ask any vendor in this category: for any given dollar of last month’s AI spend, can you tell me what it was for, and could you have spent it better?