AI Evaluation(Updated )· 14 min read

How to Calculate AI ROI Without Guessing

TG

Trupti Gavit

Founder, KaryoWorks

The ROI problem

Ask any team "is this AI tool worth it?" and you will get one of two answers:

  1. 1."Absolutely!" (no data)
  2. 2."I think so..." (no data)

AI ROI is treated like a feeling rather than a calculation. But it is math — if you have the right inputs and subtract the costs that most ROI models ignore.

The scale of the measurement gap is significant. An IBM CEO study found that only 25% of AI initiatives delivered expected ROI over the past three years, and only 16% scaled enterprise-wide. IBM's Think Circle research reports that just 29% of executives can measure AI ROI confidently — while 79% see productivity gains they cannot translate into financial impact.

Why most AI ROI calculations fail

Productivity without conversion

IBM's research notes that 79% of organizations observe productivity gains from AI — but converting those gains into financial impact requires explicit workflow design. Time saved on drafting does not become ROI until that time is redeployed to higher-value work, headcount is avoided, or output volume increases.

Pilot ROI does not scale

IBM's AI research found that as AI pilots scale to production, returns moderate to roughly 7% ROI on average — below the typical 10% cost-of-capital hurdle rate. The top decile of organizations achieves approximately 18% ROI. The difference is not better models — it is better measurement, workflow redesign, and organizational integration.

Enterprise impact vs. tool impact

McKinsey's State of AI 2025 reports that 88% of organizations use AI, but only 39% report EBIT impact at the enterprise level. Individual tool ROI can be positive while enterprise ROI remains invisible — because tools are deployed without connecting to business outcomes.

The formula

Monthly AI ROI = Total Monthly Value Generated - Monthly Cost

Where:

Total Monthly Value = Time Savings Value + Quality Improvement Value + Error Reduction Value - Rework Cost

The subtraction matters. Most ROI models ignore it.

Productivity benchmarks from published research

Before estimating time savings from gut feel, anchor your assumptions. The Alice Labs AI Automation ROI Benchmark 2026 synthesizes peer-reviewed and enterprise case data:

Use caseDocumented gainSource type
Customer support15% productivity gainPeer-reviewed (5,179 agents)
Professional writing40% faster completionControlled study (Noy & Zhang)
Coding (controlled task)55.8% faster completionGitHub Copilot experiment
Knowledge work (suitable tasks)12.2% more tasks, 25.1% fasterHBS/BCG jagged-frontier study
Copilot-style deployments1.9–4.0 hours saved per weekCross-case enterprise data

How to use these benchmarks

These are ceiling references, not guarantees. The HBS/BCG study also found 19 percentage points worse correctness outside the AI "frontier" — meaning gains on suitable tasks coexist with losses on unsuitable ones.

For ROI modeling, use the lower bound of each range for conservative estimates and apply benchmarks only to the specific task type — not the whole job. Adjust for your team's adoption rate (enterprise average: 34% per Zylo data).

Time savings value

This is usually the largest component — and the easiest to overestimate.

Formula: Users x Hours saved per week x Hourly rate x 4.33 weeks/month

Conservative benchmark: Use 1.9 hours per week (lower bound from Alice Labs) multiplied by your adoption rate.

Example at full adoption: 5 content writers save 1.9 hours/week each. Loaded hourly rate: $45.

Time Savings Value = 5 x 1.9 x $45 x 4.33 = $1,852/month

Compare this to an estimate using 3 hours/week ($2,924/month). The benchmark-backed number is 37% lower — and more defensible in a finance review.

Quality improvement value

Quality gains are real but harder to quantify. Alice Labs documents 40% faster completion on bounded professional writing with higher rated quality on average. But "faster" is not automatically "better."

Assign a dollar value only where you can quantify it. If quality "feels better," estimate $0 until measured.

Error reduction value

Some AI tools prevent mistakes with real costs: AI code review catching bugs, AI validation catching data entry errors, AI grammar tools preventing client-facing errors.

Formula: Errors caught per month x Average cost per error

Example: AI code review catches 4 bugs/month that would each take 3 hours to fix in production. Developer rate: $75/hour.

Error Reduction Value = 4 x 3 x $75 = $900/month

Only include errors you can document.

Net productivity: Subtract rework and verification time

This is the adjustment most ROI calculations miss.

AI output requires verification. Drafts need editing. Code needs review. Data needs checking. The time AI saves on generation is partially consumed by the time humans spend confirming accuracy.

Why net productivity matters

Alice Labs cites the HBS/BCG jagged-frontier finding: AI produces 19 percentage points worse correctness outside its capability frontier. If your team uses AI on tasks beyond that frontier, verification time increases — potentially erasing time savings.

How to calculate rework cost

Rework Cost = Verification hours per week x Users x Hourly rate x 4.33

AI output qualityVerification time as % of generation time saved
Minor edits only10–20% of time saved
Regular major revisions30–50% of time saved
Frequent rejection/regeneration50–80% of time saved

Example: AI saves 1.9 hours/week per writer, but verification consumes 0.5 hours/week.

Net time saved = 1.9 - 0.5 = 1.4 hours/week (26% rework tax)

Adjusted monthly value = 5 x 1.4 x $45 x 4.33 = $1,364/month (vs. $1,852 gross)

See Your AI Is Lying to You for the five most common hallucination patterns that drive rework.

For ROI purposes, treat verification time as a real cost:

  • Low-risk tasks (internal brainstorming, rough drafts): 10% rework tax
  • Medium-risk tasks (client emails, reports with data): 25% rework tax
  • High-risk tasks (published content, legal, financial): 40–50% rework tax

Putting it together: Benchmark-backed example

Tool: AI writing assistant for a 5-person content team. Monthly cost: $300

ComponentGross estimateRework adjustmentNet value
Time savings (1.9 hrs/wk, Alice Labs lower bound)$1,852-26% verification tax$1,364
Quality improvement (revision rate -20%)$180$180
Error reduction$0$0
Total monthly value$2,032$1,544
Monthly cost$300
Monthly ROI$1,244
ROI percentage415%

Even with conservative benchmarks and rework subtracted, a well-fit tool shows strong ROI. Compare to a poorly-fit tool:

ComponentValue
Time savings (5 users, 0.3 hrs/wk actual, 40% adoption)$117
Rework (high rejection rate)-$58
Quality improvement$0
Net monthly value$59
Monthly cost$300
Monthly ROI-$241 (negative)

Same formula. Different inputs. The math tells you what to cut.

When ROI is negative

If monthly cost exceeds net monthly value, check before cancelling:

  1. 1.Is adoption the problem? At 34% average utilization (Zylo), most tools underperform on adoption before they underperform on capability.
  2. 2.Is the use case wrong? A powerful tool applied to the wrong task will show usage without value.
  3. 3.Is it a timing issue? Some tools need 3–6 months for proficiency. But 44% of AI licenses are abandoned within 90 days.
  4. 4.Is rework eating the savings? Recalculate with verification time included.

Rule: Negative net ROI after 6 months with adoption above 40% = cut the tool.

The sensitivity check

Always stress-test your assumptions. Optimistic ROI models are how 75% of AI initiatives fail to deliver expected returns (IBM).

Run three scenarios:

ScenarioTime savings assumptionRework taxWhen to use
ConservativeLower bound benchmark x adoption rate30–50%Default for finance review
ExpectedMid-range benchmark x adoption rate20–30%Internal planning
OptimisticUpper bound x 100% adoption10–15%Vendor business case only

Decision rules:

  • ROI positive in conservative scenario = KEEP with confidence
  • ROI positive only in expected scenario = OPTIMIZE (improve adoption or reduce rework)
  • ROI positive only in optimistic scenario = CUT or REPLACE
  • ROI negative in all scenarios = CUT immediately

Context for what "good" looks like at scale: IBM's research puts average scaled AI ROI at 7%, with top-decile organizations at 18%. A tool-level ROI of 200–400% on a specific workflow is plausible. Enterprise-wide AI ROI claims above 18% require exceptional evidence.

Start calculating

Most businesses have never run this calculation for any AI tool. Doing it once for your three most expensive subscriptions — with benchmark-backed inputs, rework subtracted, and sensitivity tested — will tell you more than six months of "I think it's helping."

The AI Automation Audit System includes a pre-built ROI Calculator spreadsheet that automates these formulas. Or start with the free AI Spend Calculator to establish your cost baseline.

For warning signs that you need this calculation now, see 5 Signs You're Wasting Money on AI.

Sources

TG

Trupti Gavit

Founder, KaryoWorks

AI practitioner building evaluation and decision systems for businesses managing AI investments.

Get practical AI frameworks by email

Evaluation systems, spend analysis, and decision tools — one email when we publish. No spam.