Introducing AIR: A Practical Way to Measure AI ROI
Alan Hand
June 8, 2026
11 min read
In August 2025, MIT's NANDA initiative published The GenAI Divide: State of AI in Business 2025, the most-cited piece of AI-ROI research of the year. The headline finding: 95 percent of enterprise generative AI pilots produced no measurable financial return, against roughly $30 to $40 billion in enterprise AI spend. The study drew on more than 300 enterprise deployments, 52 case studies, and 153 leadership surveys.
What unsettled CFOs most was not the 95-percent number itself. It was that the same organizations were continuing to spend, often aggressively, without a defensible framework for telling the 5 percent of bets that worked from the 95 percent that did not. The CFO's question, "what are we actually getting for it," had no defensible answer at scale.
That is the gap AIR (AI Return) is built to close.
AI ROI is structurally different from cloud ROI
Traditional ROI works when inputs and outputs are clean. You spend X, you earn Y, you measure the delta. Cloud ROI in the early 2010s eventually settled into this shape after a few years of confusion: you compared cloud spend to the equivalent on-prem spend, you adjusted for capex versus opex, you got a defensible percentage. The number was hard to compute but conceptually simple.
AI ROI does not behave that way. What you get from AI is faster workflows sometimes, better outputs sometimes, more experimentation always, and rising costs always. The output is qualitative and aggregated across thousands of microscopic interactions. The cost is concentrated in vendor invoices that bundle dissimilar work. The mapping between input and output is many-to-many.
The result is a familiar trap: lots of activity, unclear impact, a budget line that finance cannot tie to anything specific. We have watched this play out at customer after customer through 2026. Teams that figure out how to measure AI value early end up with budget authority and political room to expand. Teams that cannot defend the spend get squeezed when the next cost review comes around.
The definition
AIR answers one question.
For every dollar spent on AI, how much value are we actually generating?
The definition:
AIR=[ \frac{\text{Value Created by AI} - \text{AI Cost}}{\text{AI Cost}} ]
That is the whole thing. An AIR of 1.0 means AI is paying for itself in net new value. An AIR of 3.0 means every dollar of AI spend is producing three dollars of value, net of cost. AIR below zero means you are losing money, which is a perfectly legitimate finding early in adoption. AIR of exactly 0.0 means AI is breaking even, neither earning nor costing the organization on a net basis.
The metric is intentionally simple. The work is in defining the inputs.
Making AIR real
To be useful, AIR has to be grounded in measurable components. The expanded form:
AIR=[ \frac{\text{Revenue Gain} + \text{Cost Savings} + (\text{Time Saved} \times V) - \text{AI Cost}}{\text{AI Cost}} ]
Where:
- Revenue Gain is new revenue enabled by AI. Features shipped faster that earned earlier. Products that would not have existed without AI. New customer segments served because AI made the unit economics work.
- Cost Savings is reduced labor, tooling, or infrastructure spend that AI made possible. Headcount avoided. Vendor contracts canceled because the AI tool replaced them. Lower support volume because AI deflected tickets.
- Time Saved is measurable productivity improvement, captured per task or per workflow.
- V is the value of time, the variable everything hinges on. More on this in the next section.
- AI Cost includes tokens, infrastructure, engineering integration time, evaluation work, and tooling.
The expanded form forces the conversation finance actually wants to have: What did we gain? What did we save? What is an hour of engineering time worth? What did this all cost? None of those questions are answered by a usage dashboard.
The critical variable: V
The model hinges on one question.
What is an hour of work actually worth?
Three defensible answers exist, and the right one depends on who is asking.
- Conservative, finance-friendly. Use fully loaded cost. A senior engineer at $150K base lands around $100 to $150 an hour all-in once you include benefits, equity, payroll taxes, real estate, and shared overhead. CFOs trust this number because they already use it in headcount planning. It is the easiest V to defend and produces the lowest, most conservative AIR.
- Operational, more accurate. Use value per task. A bug fix carries different weight than a new feature, which in turn carries different weight than a revenue-generating workflow. Operational AIR breaks work down by task type and applies a different V to each, anchored in either the historical cost-per-unit or the historical value-per-unit for that task category. This V is harder to assemble but produces a more honest picture of where AI is actually creating leverage.
- Strategic, what executives care about. Use business impact. What did faster time-to-market unlock? What hires were avoided? What shipped that otherwise would not have? Strategic AIR uses business outcomes as the unit of value rather than engineer-hours. It is the hardest to defend rigorously, because attribution is messy, and the most powerful in a board conversation, because it speaks the language of strategy rather than the language of engineering operations.
Most organizations should start conservative, move to operational once they have task-level data, and reach for strategic V only when the business case demands it. Mixing the three within a single AIR calculation is the fastest way to produce a number nobody trusts.
Why most AI ROI calculations fail
We have read a lot of AI ROI calculations through 2026, from internal models customers built for their own boards to vendor case studies trying to justify renewal pricing. The failures cluster into three patterns.
- Counting time saved as value. Saving an engineer two hours does nothing if the team's output stays the same. Those hours have to translate into more shipped features, fewer hires, or measurably better quality. Time saved is a signal that something might be working. By itself it does not count as a result. The companies that get this right track time saved as an input variable, then look downstream for the actual outcome change, and only count the outcome.
- Ignoring real costs. AI cost includes far more than the API bill. Engineering integration, prompt iteration, ongoing maintenance, the evaluation infrastructure needed to know whether the AI is doing what you think it is doing, tooling, and the opportunity cost of the team working on AI instead of the next thing all belong in the denominator. Calculations that count only tokens dramatically understate the cost side.
- Double-counting gains. "Faster and cheaper and better" sounds compelling in a slide, and the gains usually overlap. If AI lets you ship a feature in half the time, the time savings and the additional features shipped describe the same outcome from two different angles. Claiming both inflates the result. Disciplined AIR picks one frame, the one that maps most cleanly to the value the organization is actually trying to create, and stays with it.
A fourth pattern, less common but worth flagging: borrowing vendor case study numbers. "Vendor X claims customers save 30 percent on developer hours." Even if true on average, the variance is enormous. Your AIR has to be computed from your own data or it is not your AIR.
Task-level AIR
Treating all work as equally valuable is the single biggest mistake we see. Real AIR rolls up from the task level:
AIR=[ \frac{\sum_{i=1}^{n} (\text{Time Saved}_i \times \text{Value}_i) - \text{AI Cost}}{\text{AI Cost}} ]
Each task type carries its own time savings and its own value multiplier. A few categories worth tracking separately:
- Bug fix. Value tied to the cost of the bug had it shipped. High variance: a bug that would have caused an outage is worth orders of magnitude more than a cosmetic defect.
- Feature ship. Value tied to time-to-market acceleration. Highly dependent on whether the feature was on the critical path of a revenue-generating launch.
- Documentation. Value tied to onboarding velocity and reduced support load. Lower per-task value but large aggregate volume.
- Code review. Value tied to throughput. Indirect impact on cycle time, which compounds across the organization.
- Customer support response drafting. Value tied to time-to-resolution and CSAT. High per-organization volume with substantial aggregate effect.
This is where AIR becomes defensible. The claim shifts from "AI made us 20 percent more productive" to a per-task computation grounded in actual time-savings data from your own systems. The first claim is unfalsifiable. The second can be audited.
What AIR makes possible
Once you measure AIR properly, several questions that have frustrated engineering and finance leaders for two years become answerable.
- Comparing tools. Is Copilot earning its keep? Is the new agent framework actually helping? AIR turns the question into a number computed from your own data rather than the vendor's. Two tools tracked in parallel for a quarter produce comparable AIR figures, and the better tool wins on merit.
- Optimizing usage. Which prompts and models are wasteful? Where are tokens being burned with no return? The lowest-AIR workflows surface first and become the obvious targets for caching, model downshifts, or process redesign.
- Justifying spend. When finance asks why you are spending $50K a month on AI, the answer becomes, "because it is generating 3.2x AIR on a $50K investment. Here is the per-team breakdown. Show me a better use of that capital and I will switch." This is a conversation finance can have. The vague "AI is helping us ship faster" is not.
- Driving engineering decisions. Where should AI be applied? Where should it not be? AIR makes "should we" a quantitative question rather than a stylistic one. Teams that adopted AI on faith and ended up with sub-1.0 AIR for specific workflows can defund those workflows without losing face, because the data made the call for them.
From cost to performance
Most organizations treat AI today the way they treated cloud in 2014: as a cost center with fuzzy upside. The teams that won the cloud era spent years turning cloud spend into a measurable performance layer through FinOps. The discipline did not exist as a named practice in 2014. By 2020 it was a category. By 2024 it was a board-level conversation.
AIR is the same move applied to AI. The pattern is familiar: a new infrastructure cost category becomes legible once a defensible headline metric emerges. FinOps got there with the Effective Savings Rate. DevOps got there with DORA's lead time and deployment frequency. AIR is the candidate headline metric for AI.
The headline-metric move is not cosmetic. A discipline without a defensible headline metric stays academic. A discipline with one shows up in board decks, gets attached to executive bonuses, and accumulates institutional muscle.
We have built AIR into TokenIQ, the AI pillar of the Xosphere IQ platform launched today. The four TokenIQ phases are designed to maximize AIR: Phase 1 attributes cost to business outcomes, Phase 2 estimates value, Phase 3 computes AIR itself, and Phase 4 closes the loop with optimization.
If you want to see how the four phases surface AIR for your own AI workloads, reach out. We are working with a small group of design partners through the rest of 2026 ahead of broader availability.