A Point of View

The Coming Cost Reckoning for Enterprise AI

Establishing what enterprise AI costs to operate, and who answers for that cost.

The price is not the cost

A published rate is a starting figure. What an institution actually spends to put a model into trustworthy production is a different number, and the distance between them is the subject of this paper.

A published rate tells the buyer what the provider charges for a defined unit of service. Under pure consumption, the provider carries the utilization risk of the capacity behind that service. The institution still has to account for the operating costs and control obligations that remain on its own side of the transaction.

For an institution operating under supervision, an output has to meet the control requirements of its intended use. The work required to explain, monitor, retain, and audit it belongs in the production cost alongside the provider's rate.

This is the distinction the rest of the paper builds on. An institution that treats a published rate as its complete operating-cost model leaves out the costs and obligations it retains.

Treating intelligence this way, as a managed input with its own cost structure, utilization behavior, and accountability, is the beginning of a discipline this paper calls intelligence economics. Intelligence economics emerges now because machine intelligence has become a factor of production large enough to require its own economic discipline. The scale is no longer in question. Gartner forecast generative AI spending of roughly $644 billion for 2025, and in January 2026 forecast total worldwide AI spending, a broader category, at approximately $2.5 trillion for the year. At that scale, the absence of a discipline to govern the spending is itself the problem. What follows is a first account of that discipline, built around a method for constructing the real cost and a model for governing it.

Conceptual flow comparing pure consumption, spending commitments, reserved capacity, and ownership before governance costs are considered. Pure consumption requires no separate buyer-side adjustment for provider idle capacity or asset life. Spending commitments and reservations create payment obligations. Ownership allocates capital cost over the asset's service life. Governance costs apply across all routes.
Exhibit 1. From quoted rate to loaded cost. The adjustments depend on how capacity is bought. Under pure consumption, the buyer does not separately adjust the usage rate for the provider's idle capacity or asset life. Commitments oblige the buyer to meet a spending minimum or pay for reserved capacity. Ownership spreads capital cost across an asset's service life. Governance costs apply across all three.
Open full size

We have seen this before

Public cloud went through the same arc: capability first, cost crisis second, governance third. AI inference is entering the second phase now.

The trajectory is familiar because the industry lived it with public cloud. Adoption began as a capability question, asking whether the workload could run there at all and what that unlocked. Cost was a secondary concern, tolerated as the price of speed and elasticity. Then the invoices arrived, and they were large, opaque, and difficult to attribute. Nobody could say with confidence which team, product, or decision had generated which portion of the bill. The spend was real, but it was unowned.

That cost crisis produced a discipline. It went by the name FinOps, but the substance was older than the label: make the spend visible, make it attributable, make it someone's responsibility. Tagging, cost allocation, showback, chargeback, unit economics, reserved-instance planning. None of it was glamorous, and all of it became table stakes for running cloud at scale in a serious institution.

The parallel has since stopped being an analogy. In August 2026 the Linux Foundation launched the Tokenomics Foundation, a vendor-neutral body chartered to set standards, benchmarks, and best practices for the economics of AI, with roughly thirty founding members including JPMorgan Chase, BNY, IBM, Accenture, Oracle, SAP, and ServiceNow. It operates in partnership with the FinOps Foundation, and one executive director leads both. Its published roadmap covers value metrics for AI return, vendor-neutral models for the full cost of AI, standard methods for measuring cost to serve, token cost telemetry carried in the FOCUS billing specification, and practitioner certification.

Two of the founding members are banks. That is the detail worth pausing on, because it inverts the assumption that supervised institutions arrive late to a cost discipline. They are among the parties defining this one, for the same reason they built model risk management before it was general practice. Where an output has to be explained to a supervisor, the accounting for what that output cost is already half constructed.

Inference now sits where cloud sat at the start of its cost crisis, and the early signals are already visible. The FinOps Foundation's annual survey found the share of surveyed practitioners' organizations managing AI spend rising from under a third to roughly two-thirds in a single year, and to nearly universal the year after. Forecasts now put inference ahead of training in AI-optimized infrastructure services for 2026, the marker of a technology moving from pilots into production at scale. Yet attribution lags badly: few institutions can answer the basic question of which business activity consumed which portion of inference cost, and whether it was worth it. The institutions that answered that question early for cloud spent less and governed better than those that waited. The same opportunity is open now, and for the same reason: the discipline can be built before the crisis forces it.

The parallel instructs, but it is not exact, and the differences are where a borrowed method fails. Inference cost is far more volatile than compute cost: it can swing by an order of magnitude week to week, driven by model choice, prompt design, and output length in ways raw infrastructure never was. Model selection is a cost lever with no clean cloud precedent, since the same task routed to a smaller model can cost a fraction of the larger one, trading against quality in a way provisioning decisions never did. A single poorly formed prompt can cost many times an optimized one, which means engineering choices that look purely technical are now financial ones. And the work of making an output defensible attaches to the output and the decision, not to a machine, so it resists the infrastructure-style allocation that cloud practice relies on. A method lifted whole from cloud cost management will miss all of this. One informed by it and adjusted for these differences will not.

There is a deeper difference, and it cuts in the institution's favor rather than against the thesis. Cloud governed infrastructure: machines, storage, network, things one step removed from any business decision. Inference governs the output itself, which feeds directly into decisions that carry regulatory and reputational consequence. The object under management is no longer a resource but a judgment. That raises the stakes of attribution rather than lowering them, because a misallocated compute cost is an accounting error while a misgoverned inference is a decision the institution may have to defend. Inference, in other words, may demand more governance than cloud ever did, not less, precisely because what it produces is closer to the business.

A skeptic will argue that model prices are falling too quickly to justify building a discipline around the cost. That argument treats a falling unit price as a falling operating bill. Inference unit cost has fallen sharply. Stanford's AI Index tracked the price of inference at a fixed performance level falling from around twenty dollars to roughly seven cents per million tokens over about two years. Over the same period total enterprise AI spend rose anyway, because usage expanded faster than price fell. A price drop removes friction that was holding usage down, and where consumption expands faster than the price falls, the bill grows rather than shrinks. More important, the governance load does not fall on the same curve as the model. Controls, audit, explainability, and oversight are tied to risk and regulation, not to the price of compute, so as the model gets cheaper the control environment becomes a larger share of total cost, not a smaller one. Falling model prices make governance more of the problem, not less. Goldman Sachs Research forecasts a twenty-four-fold increase in token consumption between 2026 and 2030. The effect on spending depends on how prices and the mix of usage change over that period.

June 2026 added a complication the falling-price assumption cannot absorb. The most capable generally available model launched at double the per-token rate of the flagship it replaced.

the discipline can be built before the crisis forces it

Exhibit 2: the same arc, one cycle apart
Exhibit 2. The same arc, one cycle apart. Cloud has traversed all three phases; inference is entering the second.

What a standard settles, and what it does not

A common standard can settle the arithmetic of inference cost. What that arithmetic is worth inside one institution, and whether the output can be defended, is decided somewhere no standard reaches.

A standards body normalizes measurement, and that is the layer this discipline has been missing. Telemetry carried in a common billing specification makes inference spend legible across providers, which turns the instrumentation described below from a project into a specification an institution can adopt. That is a real contribution and it should be taken up rather than duplicated.

What a vendor-neutral standard cannot settle is what the measurement means inside a particular institution. Thirty organizations spanning hyperscalers, resellers, software vendors, and banks can agree on how a token is counted. They are unlikely to agree on what it costs to make that token's output defensible under supervision, because that figure is produced by a control environment none of them share. Governance load is contested and sector-specific. In regulated institutions it is also the layer that moves the total. A common way to count is what the standard can produce. How much control an institution buys, when in the build it buys it, and what it will accept as evidence that the spend returned something remain institutional judgments.

The relationship is a stack rather than a rivalry. The standard settles the unit, the telemetry, and the basis for comparing providers. The discipline decides which capacity model fits the observed utilization, how wide the governance layer is permitted to become, whose budget carries the total, and what the institution received in exchange. A standard tells an institution what it spent. It does not tell the institution whether it should have.

Exhibit 3: three layers of the same number, set by the provider, the standard, and the institution
Exhibit 3. Three layers of the same number. The provider sets the quoted rate, the standard sets the unit and the telemetry, and only the top layer is an institutional judgment.

Visibility, reduction, accountability

A cost-governance practice for inference runs in phases. Attribution requires a cost total and a defensible basis for assigning it. Direct measurement can support that basis, while shared costs may require agreed allocation rules. The later ordering is a sequencing recommendation, and its reasons are organizational rather than technical.

Phase one. Visibility. Instrumentation builds a current picture of spending and the workloads behind it: what is being spent, on which models, by which workloads, at what utilization, and under which capacity model. This is the inference equivalent of cloud tagging and cost allocation. The deliverable is a decomposition of inference cost into its real components, including the governance load the vendor invoice never shows. Reduction is a different case. An institution can stop unnecessary work before detailed instrumentation exists. A baseline and subsequent measurement let it quantify the saving, assess the effect on quality and latency, and track whether the result holds.

Phase two. Reduction. Cost visibility helps identify potential savings. Each change still has to be assessed against the workload's quality, latency and control requirements. Is a workload on the capacity model that fits its utilization, or is it paying consumption rates for steady, predictable demand? Is it routed to a more capable and more expensive model than the task requires? Automated routing can lower direct inference cost. It does not by itself establish how the institution allocates that cost, who answers for it, or whether the controls meet its obligations. Is reserved capacity sitting idle, converting a discount into a loss? Is governance load being counted at all? Early cost reductions can help build support for changes to allocation and budget responsibility.

Phase three. Accountability. Chargeback assigns inference cost to consuming units and makes it a budget responsibility. The allocation method must be defensible, executives must support it, and the affected units need a way to challenge how their share was assigned. Chargeback is a change-management problem wearing a cost-model costume.

An institution can establish budget responsibility without internal chargeback when an accountable owner holds and exercises authority over the workload's budget. The owner must be able to act on the spending. A claimed return still requires evidence against an agreed baseline.

Chargeback is a change-management problem wearing a cost-model costume.

Exhibit 4: three phases, visibility, reduction and accountability, joined by dashed links showing support rather than a required input. A panel beside the accountability card carries the one requirement, a cost total and a defensible allocation basis. A dashed loop returns a budget consequence to visibility.
Exhibit 4. What each phase hands the next. Visibility supports reduction, and early reductions can support the move to accountability, neither as a precondition. Attributing cost is the one step with a requirement of its own.
Open full size

The Loaded Cost Framework

To manage the economics of intelligence consistently, an institution needs a consistent way to construct cost. The Loaded Cost Framework does this in four layers.

The published rate is a single number standing in for a calculation the institution should be doing itself. The Loaded Cost Framework is that calculation, built up the same way every time so that figures from different teams rest on the same basis and can be compared. It operates at a deliberately high level here; the detailed mechanics are established in engagement. The structure, four layers, is stable across them.

Direct inference cost. The measurable spend to produce output: tokens consumed, the model used, and the rate paid under the operative capacity model. This is the layer vendors price and the only one most institutions currently see. Within a single model, the effort setting has become a priced dimension of its own. Effort settings affect the cost of finished work and can be adjusted task by task.
Capacity-model adjustment. The same workload carries a different effective cost depending on how capacity is bought, and the buyer's exposure to underuse begins when the buyer commits. A capacity commitment reserves capacity that the buyer pays for whether or not it is fully used. A spending commitment promises an amount of spend over a term in exchange for a discounted rate, and any shortfall is assessed over the contract's applicable commitment period. Both create exposure, and they do not put the buyer on the hook for the same thing. Consumption is a variable cost that scales with use. Reserved capacity is a lower rate paid in exchange for carrying the risk of capacity left unused. Ownership spreads capital and fixed operating costs over output, making utilization central to its economics. Whether it costs less than consumption or reserved capacity depends on the institution's capacity, amortization, and operating assumptions. The model chosen, set against actual utilization, can move effective cost severalfold. The interaction across teams matters as much as the model itself. A temporary test workload can distort the demand history used to size the next commitment. If planners treat that peak as continuing demand, the institution can be left paying for capacity after the test has ended. Capacity decisions made in isolation tend to cost more than the same decisions coordinated. Jurisdiction has also become a capacity variable. Anthropic publishes a premium for US-only inference on specified models and platforms. It is a surcharge whose justification belongs to risk and compliance while its cost lands in a technology budget, which makes it an allocation question from the day it is incurred.
Governance load. The cost of making an output defensible under supervision: model risk management, explainability, monitoring, audit, and data control. This layer is largely absent from the vendor invoice, and its size is not fixed. Planning controls before deployment can reduce avoidable rework and retrofit costs. Timing is one variable the institution can influence, alongside use-case risk and requirements that change during the build. This layer is no longer entirely the institution's to schedule. Provider terms can make governance an adoption decision before implementation begins. Retention requirements and the conditions attached to exceptions affect whether an institution can use a model for a particular workload. Safety fallback can also divide a workflow across models at different rates. The institution has to establish which data-handling terms apply and account for the billable work each model performs.
Allocation basis. The rule by which shared and committed costs are attributed to consuming units. Reserved and owned capacity in particular are shared resources whose cost must be divided on a basis the organization will accept as fair. The allocation basis is the bridge from a cost model to a chargeback model, and the point at which the analytical work meets the political work.

Put simply: the direct cost is the figure everyone sees, the capacity-model adjustment determines what the effective rate actually becomes at real utilization, the governance load is the cost of meeting the institution's control obligations, and the allocation basis decides whose budget ultimately carries the total. A practice that keeps all four in view produces a number that withstands examination. One that collapses them back into a single quoted rate produces a number that does not.

Beneath the allocation basis sits a prior question the framework must now ask explicitly: what unit the cost is denominated in. The structures most institutions inherited assume the unit of consumption is an interaction, a seat, a query, a call. The current model generation undermines that assumption. A model that works autonomously across an extended run returns a finished piece of work rather than a stream of answers, and the deliverable becomes the unit the economics organize around. Published CursorBench results report average cost per task alongside benchmark scores.

For institutional costing, the measure is cost per accepted task. The numerator includes the costs of failed attempts, retries and required human review across the workload being measured. The denominator counts completed tasks that met a defined quality threshold. Counting failed attempts as additional accepted work would make an inefficient system look cheaper. The usage reported for a returned answer may omit earlier billable attempts required to produce it. Measuring the task means accounting for those attempts as well.

When the unit of value is a finished task, per-seat and per-query allocation misprice in both directions at once: exploratory use is overcharged relative to the value it creates, while a single long-running task that consumed what a thousand interactions would is buried inside an averaged rate no one examines. The allocation basis decides whose budget carries the cost. The unit of account decides what is being counted, and it is moving from the interaction to the deliverable.

Exhibit 5: capacity cost per normalized unit of inference output for consumption, reserved and ownership capacity. Reserved capacity and consumption are equal at 55% sustained utilization, and reserved is cheaper above it, on illustrative assumptions.
Exhibit 5. The Loaded Cost Framework: capacity cost per normalized unit of inference output across the three capacity models, on illustrative assumptions. Reserved capacity and consumption are equal at the annotated 55% boundary; reserved is cheaper above it.
Open full size

The Accountability Ladder

Institutions do not arrive at accountability in one step. They climb to it, and the rungs are defined by how much of the cost anyone actually answers for.

The Accountability Ladder distinguishes visibility, allocation, budget responsibility and defensible return. It locates gaps in accountability. Institutions can establish budget responsibility through chargeback or through an accountable owner who holds and exercises authority over the workload's budget.

Opaque. Inference spend exists but is neither measured nor attributed. The bill is paid centrally and questioned only when it grows alarming. Most institutions begin here.
Visible. Spend is instrumented and decomposed. Leadership can see what is spent and on what, but cost remains a central line item that no consuming unit feels.
Showback. Cost is attributed to consuming units for visibility, without yet moving budgets. Behavior begins to shift through awareness alone, and the allocation method is tested in lower-stakes conditions before money rides on it.
Chargeback. Attributed cost moves real budgets. Consuming units own their inference spend, and consumption decisions carry the discipline that ownership imposes. The rungs describe increasing strength of consequence rather than organizational maturity.
Return. Attributed cost is set against what the consuming unit received for it, on a basis the institution will defend. This is the rung the discipline is furthest from, and the one supervised institutions are best positioned to reach, because part of what a defensible output returns is denominated in a currency they already track: remediation not required, findings not raised, restrictions not imposed. An institution at chargeback knows whose budget carries the cost. An institution at return knows whether the cost was worth carrying.

A caution belongs with the return rung, because it is the one most easily faked. Attributing value invites the same optimism that produced a decade of unfalsifiable cloud savings claims, and the guard is the one that already gates the rungs below it. The basis on which return is claimed has to be accepted by the paying unit before the number is published, not defended after. A stated baseline belongs with it. Absence of regulatory findings does not by itself establish avoided cost, because that is a claim about what would otherwise have happened, and the comparison has to be named along with the assumptions it rests on.

Showback gives consuming units an opportunity to examine and challenge the allocation before it changes their budgets. For return, the corresponding requirement is an agreed basis for attributing cost and assessing the outcome.

Exhibit 6: five rungs from opaque to return with the requirement for each transition. Two routes reach return, one through chargeback and one bypassing it through an accountable owner who holds and exercises the budget. The return card states the requirement common to both.
Exhibit 6. The Accountability Ladder: opaque to return, with what each rung asks of the one below it. Return can be reached through chargeback or through an accountable owner who holds and exercises the workload's budget, and either route requires attributed cost and evidence of outcomes.
Open full size

What to build before the reckoning

The argument reduces to a few commitments an institution can begin now, ahead of the pressure that will eventually compel them.

  1. Treat the published rate as a figure to interrogate, not a number to budget. Separate out the assumptions behind it before it anchors any decision.
  2. Establish the cost total and a defensible allocation basis. Use instrumentation to improve visibility and evaluate savings against a stated baseline.
  3. Adopt one reference cost model across the institution, governance load included, so figures from different teams can be compared and defended on the same basis.
  4. Plan controls before deployment. Early planning can reduce avoidable rework and retrofit costs.
  5. Approach chargeback through showback. Showback gives consuming units an opportunity to examine and challenge the allocation before it changes their budgets.
  6. Adopt the emerging measurement standard rather than building a parallel one, and spend the saved effort on the layers the standard does not reach: governance load, allocation basis, and the return claimed against them.

What counts as success is easily mistaken for spending less. An institution that cut its inference bill by routing every workload to the cheapest model and letting its controls lapse would have lower costs and worse economics: spend it still could not attribute, on outputs it could no longer defend. Success is inference spend that someone owns and can account for. Lower spend may follow, but it is a result, not the objective.

Notes on sources

Figures cited indicate scale and direction and are drawn from named industry sources. Generative-AI spending of roughly $644 billion for 2025 is Gartner's figure (forecast March 2025, a 76 percent rise over 2024); total worldwide AI spending of approximately $2.5 trillion for 2026 is Gartner's January 2026 forecast. The rise in organizations managing AI spend, from roughly 31 percent to 63 percent across 2024-2025 and to near-universal in the 2026 survey, is from the FinOps Foundation's annual State of FinOps survey. The crossover of inference past training in AI-optimized infrastructure services spending reflects 2025-2026 forecasts from Gartner and Deloitte. The decline in inference unit cost, from roughly twenty dollars to about seven cents per million tokens over approximately two years, is from Stanford's AI Index 2025.

The Tokenomics Foundation was launched by the Linux Foundation on August 4, 2026, following a June 3, 2026 announcement of intent; founding membership, roadmap scope, and the partnership with the FinOps Foundation are from the Linux Foundation's launch materials. Founding members named in the text are a partial list of approximately thirty.

Goldman Sachs Research forecasts a twenty-four-fold increase in token consumption between 2026 and 2030. The forecast appears in its article ‘AI Agents Forecast to Boost Tech Cash Flow as Usage Soars,’ published May 20, 2026.

The distinction between a capacity commitment and a spending commitment follows the provider's own description of the two structures; see the AWS Savings Plans FAQ. On assigning shared costs through documented allocation rules rather than direct metering at the eventual allocation level, see the FinOps Foundation's allocation guidance.

Cursor's Spring 2026 Developer Habits Report presents coding benchmark scores alongside average task costs. It is cited here for that reporting approach, not as evidence of an institution's cost per accepted task. Accessed September 13, 2026.

The June 2026 launch example is from Anthropic's announcement ‘Claude Fable 5 and Claude Mythos 5,’ published June 9, 2026. It supports the per-token rate at launch, the retention policy announced with it, and safety fallback, and it is cited as a record of what was announced on that date rather than of terms now in force. The residency premium rests on the current documentation cited below.

Provider terms below were verified against primary documentation on September 13, 2026, which is an access date rather than a publication date. Covered Models carry additional retention requirements by default; the temporary zero-data-retention arrangement is limited to eligible customers and to internal business applications, can be modified or withdrawn, and leaves the Usage Policy and real-time safety controls in force (Covered Models). That internal-business-application limit attaches to the temporary arrangement rather than to every Enterprise Frontier Safeguards configuration. Enterprise Frontier Safeguards were announced on September 1, 2026 with phased rollout planned for the autumn, which is evidence of an announcement rather than of deployment to every eligible customer (announcement). Fallback billing records can distinguish earlier billable attempts from the usage reported for the returned answer (refusals and fallback). The US-only premium is a 1.1x multiplier applied across token pricing categories, including input, output, cache writes and cache reads, for Claude 4.6 and later on the Claude API first-party, Claude Platform on AWS, and Microsoft Foundry US Data Zone Standard arrangements; Bedrock and Google Cloud price regionally on their own terms (data residency).

Cost-behavior curves in the exhibits are normalized and illustrative of mechanism, not vendor benchmarks. Governance-cost effects are described qualitatively because published estimates vary widely by sector and implementation. Automated model routing is available in hyperscaler and independent platforms as of 2026; the phase-two reference above reflects that availability and makes no claim about any specific product's performance.

© 2026 Aaron Cheiffetz. All rights reserved.