A Point of View
The Coming Cost Reckoning for Enterprise AI
Establishing what enterprise AI costs to operate, and who answers for that cost.
The price is not the cost
A published rate is a starting figure. What an institution actually spends to put a model into trustworthy production is a different number, and the distance between them is the subject of this paper.
A published rate tells the buyer what the provider charges for a defined unit of service. Under pure consumption, the provider carries the utilization risk of the capacity behind that service. The institution still has to account for the operating costs and control obligations that remain on its own side of the transaction.
For an institution operating under supervision, an output has to meet the control requirements of its intended use. The work required to explain, monitor, retain, and audit it belongs in the production cost alongside the provider's rate.
This is the distinction the rest of the paper builds on. An institution that treats a published rate as its complete operating-cost model leaves out the costs and obligations it retains.
Treating intelligence this way, as a managed input with its own cost structure, utilization behavior, and accountability, is the beginning of a discipline this paper calls intelligence economics. Intelligence economics emerges now because machine intelligence has become a factor of production large enough to require its own economic discipline. The scale is no longer in question. Gartner forecast generative AI spending of roughly $644 billion for 2025, and in January 2026 forecast total worldwide AI spending, a broader category, at approximately $2.5 trillion for the year. At that scale, the absence of a discipline to govern the spending is itself the problem. What follows is a first account of that discipline, built around a method for constructing the real cost and a model for governing it.

We have seen this before
Public cloud went through the same arc: capability first, cost crisis second, governance third. AI inference is entering the second phase now.
The trajectory is familiar because the industry lived it with public cloud. Adoption began as a capability question, asking whether the workload could run there at all and what that unlocked. Cost was a secondary concern, tolerated as the price of speed and elasticity. Then the invoices arrived, and they were large, opaque, and difficult to attribute. Nobody could say with confidence which team, product, or decision had generated which portion of the bill. The spend was real, but it was unowned.
That cost crisis produced a discipline. It went by the name FinOps, but the substance was older than the label: make the spend visible, make it attributable, make it someone's responsibility. Tagging, cost allocation, showback, chargeback, unit economics, reserved-instance planning. None of it was glamorous, and all of it became table stakes for running cloud at scale in a serious institution.
The parallel has since stopped being an analogy. In August 2026 the Linux Foundation launched the Tokenomics Foundation, a vendor-neutral body chartered to set standards, benchmarks, and best practices for the economics of AI, with roughly thirty founding members including JPMorgan Chase, BNY, IBM, Accenture, Oracle, SAP, and ServiceNow. It operates in partnership with the FinOps Foundation, and one executive director leads both. Its published roadmap covers value metrics for AI return, vendor-neutral models for the full cost of AI, standard methods for measuring cost to serve, token cost telemetry carried in the FOCUS billing specification, and practitioner certification.
Two of the founding members are banks. That is the detail worth pausing on, because it inverts the assumption that supervised institutions arrive late to a cost discipline. They are among the parties defining this one, for the same reason they built model risk management before it was general practice. Where an output has to be explained to a supervisor, the accounting for what that output cost is already half constructed.
Inference now sits where cloud sat at the start of its cost crisis, and the early signals are already visible. The FinOps Foundation's annual survey found the share of surveyed practitioners' organizations managing AI spend rising from under a third to roughly two-thirds in a single year, and to nearly universal the year after. Forecasts now put inference ahead of training in AI-optimized infrastructure services for 2026, the marker of a technology moving from pilots into production at scale. Yet attribution lags badly: few institutions can answer the basic question of which business activity consumed which portion of inference cost, and whether it was worth it. The institutions that answered that question early for cloud spent less and governed better than those that waited. The same opportunity is open now, and for the same reason: the discipline can be built before the crisis forces it.
The parallel instructs, but it is not exact, and the differences are where a borrowed method fails. Inference cost is far more volatile than compute cost: it can swing by an order of magnitude week to week, driven by model choice, prompt design, and output length in ways raw infrastructure never was. Model selection is a cost lever with no clean cloud precedent, since the same task routed to a smaller model can cost a fraction of the larger one, trading against quality in a way provisioning decisions never did. A single poorly formed prompt can cost many times an optimized one, which means engineering choices that look purely technical are now financial ones. And the work of making an output defensible attaches to the output and the decision, not to a machine, so it resists the infrastructure-style allocation that cloud practice relies on. A method lifted whole from cloud cost management will miss all of this. One informed by it and adjusted for these differences will not.
There is a deeper difference, and it cuts in the institution's favor rather than against the thesis. Cloud governed infrastructure: machines, storage, network, things one step removed from any business decision. Inference governs the output itself, which feeds directly into decisions that carry regulatory and reputational consequence. The object under management is no longer a resource but a judgment. That raises the stakes of attribution rather than lowering them, because a misallocated compute cost is an accounting error while a misgoverned inference is a decision the institution may have to defend. Inference, in other words, may demand more governance than cloud ever did, not less, precisely because what it produces is closer to the business.
A skeptic will argue that model prices are falling too quickly to justify building a discipline around the cost. That argument treats a falling unit price as a falling operating bill. Inference unit cost has fallen sharply. Stanford's AI Index tracked the price of inference at a fixed performance level falling from around twenty dollars to roughly seven cents per million tokens over about two years. Over the same period total enterprise AI spend rose anyway, because usage expanded faster than price fell. A price drop removes friction that was holding usage down, and where consumption expands faster than the price falls, the bill grows rather than shrinks. More important, the governance load does not fall on the same curve as the model. Controls, audit, explainability, and oversight are tied to risk and regulation, not to the price of compute, so as the model gets cheaper the control environment becomes a larger share of total cost, not a smaller one. Falling model prices make governance more of the problem, not less. Goldman Sachs Research forecasts a twenty-four-fold increase in token consumption between 2026 and 2030. The effect on spending depends on how prices and the mix of usage change over that period.
June 2026 added a complication the falling-price assumption cannot absorb. The most capable generally available model launched at double the per-token rate of the flagship it replaced.
the discipline can be built before the crisis forces it

What a standard settles, and what it does not
A common standard can settle the arithmetic of inference cost. What that arithmetic is worth inside one institution, and whether the output can be defended, is decided somewhere no standard reaches.
A standards body normalizes measurement, and that is the layer this discipline has been missing. Telemetry carried in a common billing specification makes inference spend legible across providers, which turns the instrumentation described below from a project into a specification an institution can adopt. That is a real contribution and it should be taken up rather than duplicated.
What a vendor-neutral standard cannot settle is what the measurement means inside a particular institution. Thirty organizations spanning hyperscalers, resellers, software vendors, and banks can agree on how a token is counted. They are unlikely to agree on what it costs to make that token's output defensible under supervision, because that figure is produced by a control environment none of them share. Governance load is contested and sector-specific. In regulated institutions it is also the layer that moves the total. A common way to count is what the standard can produce. How much control an institution buys, when in the build it buys it, and what it will accept as evidence that the spend returned something remain institutional judgments.
The relationship is a stack rather than a rivalry. The standard settles the unit, the telemetry, and the basis for comparing providers. The discipline decides which capacity model fits the observed utilization, how wide the governance layer is permitted to become, whose budget carries the total, and what the institution received in exchange. A standard tells an institution what it spent. It does not tell the institution whether it should have.

Visibility, reduction, accountability
A cost-governance practice for inference runs in phases. Attribution requires a cost total and a defensible basis for assigning it. Direct measurement can support that basis, while shared costs may require agreed allocation rules. The later ordering is a sequencing recommendation, and its reasons are organizational rather than technical.
Phase one. Visibility. Instrumentation builds a current picture of spending and the workloads behind it: what is being spent, on which models, by which workloads, at what utilization, and under which capacity model. This is the inference equivalent of cloud tagging and cost allocation. The deliverable is a decomposition of inference cost into its real components, including the governance load the vendor invoice never shows. Reduction is a different case. An institution can stop unnecessary work before detailed instrumentation exists. A baseline and subsequent measurement let it quantify the saving, assess the effect on quality and latency, and track whether the result holds.
Phase two. Reduction. Cost visibility helps identify potential savings. Each change still has to be assessed against the workload's quality, latency and control requirements. Is a workload on the capacity model that fits its utilization, or is it paying consumption rates for steady, predictable demand? Is it routed to a more capable and more expensive model than the task requires? Automated routing can lower direct inference cost. It does not by itself establish how the institution allocates that cost, who answers for it, or whether the controls meet its obligations. Is reserved capacity sitting idle, converting a discount into a loss? Is governance load being counted at all? Early cost reductions can help build support for changes to allocation and budget responsibility.
Phase three. Accountability. Chargeback assigns inference cost to consuming units and makes it a budget responsibility. The allocation method must be defensible, executives must support it, and the affected units need a way to challenge how their share was assigned. Chargeback is a change-management problem wearing a cost-model costume.
An institution can establish budget responsibility without internal chargeback when an accountable owner holds and exercises authority over the workload's budget. The owner must be able to act on the spending. A claimed return still requires evidence against an agreed baseline.
Chargeback is a change-management problem wearing a cost-model costume.

The Loaded Cost Framework
To manage the economics of intelligence consistently, an institution needs a consistent way to construct cost. The Loaded Cost Framework does this in four layers.
The published rate is a single number standing in for a calculation the institution should be doing itself. The Loaded Cost Framework is that calculation, built up the same way every time so that figures from different teams rest on the same basis and can be compared. It operates at a deliberately high level here; the detailed mechanics are established in engagement. The structure, four layers, is stable across them.
Put simply: the direct cost is the figure everyone sees, the capacity-model adjustment determines what the effective rate actually becomes at real utilization, the governance load is the cost of meeting the institution's control obligations, and the allocation basis decides whose budget ultimately carries the total. A practice that keeps all four in view produces a number that withstands examination. One that collapses them back into a single quoted rate produces a number that does not.
Beneath the allocation basis sits a prior question the framework must now ask explicitly: what unit the cost is denominated in. The structures most institutions inherited assume the unit of consumption is an interaction, a seat, a query, a call. The current model generation undermines that assumption. A model that works autonomously across an extended run returns a finished piece of work rather than a stream of answers, and the deliverable becomes the unit the economics organize around. Published CursorBench results report average cost per task alongside benchmark scores.
For institutional costing, the measure is cost per accepted task. The numerator includes the costs of failed attempts, retries and required human review across the workload being measured. The denominator counts completed tasks that met a defined quality threshold. Counting failed attempts as additional accepted work would make an inefficient system look cheaper. The usage reported for a returned answer may omit earlier billable attempts required to produce it. Measuring the task means accounting for those attempts as well.
When the unit of value is a finished task, per-seat and per-query allocation misprice in both directions at once: exploratory use is overcharged relative to the value it creates, while a single long-running task that consumed what a thousand interactions would is buried inside an averaged rate no one examines. The allocation basis decides whose budget carries the cost. The unit of account decides what is being counted, and it is moving from the interaction to the deliverable.

The Accountability Ladder
Institutions do not arrive at accountability in one step. They climb to it, and the rungs are defined by how much of the cost anyone actually answers for.
The Accountability Ladder distinguishes visibility, allocation, budget responsibility and defensible return. It locates gaps in accountability. Institutions can establish budget responsibility through chargeback or through an accountable owner who holds and exercises authority over the workload's budget.
A caution belongs with the return rung, because it is the one most easily faked. Attributing value invites the same optimism that produced a decade of unfalsifiable cloud savings claims, and the guard is the one that already gates the rungs below it. The basis on which return is claimed has to be accepted by the paying unit before the number is published, not defended after. A stated baseline belongs with it. Absence of regulatory findings does not by itself establish avoided cost, because that is a claim about what would otherwise have happened, and the comparison has to be named along with the assumptions it rests on.
Showback gives consuming units an opportunity to examine and challenge the allocation before it changes their budgets. For return, the corresponding requirement is an agreed basis for attributing cost and assessing the outcome.

What to build before the reckoning
The argument reduces to a few commitments an institution can begin now, ahead of the pressure that will eventually compel them.
- Treat the published rate as a figure to interrogate, not a number to budget. Separate out the assumptions behind it before it anchors any decision.
- Establish the cost total and a defensible allocation basis. Use instrumentation to improve visibility and evaluate savings against a stated baseline.
- Adopt one reference cost model across the institution, governance load included, so figures from different teams can be compared and defended on the same basis.
- Plan controls before deployment. Early planning can reduce avoidable rework and retrofit costs.
- Approach chargeback through showback. Showback gives consuming units an opportunity to examine and challenge the allocation before it changes their budgets.
- Adopt the emerging measurement standard rather than building a parallel one, and spend the saved effort on the layers the standard does not reach: governance load, allocation basis, and the return claimed against them.
What counts as success is easily mistaken for spending less. An institution that cut its inference bill by routing every workload to the cheapest model and letting its controls lapse would have lower costs and worse economics: spend it still could not attribute, on outputs it could no longer defend. Success is inference spend that someone owns and can account for. Lower spend may follow, but it is a result, not the objective.