How to Calculate AI ROI Without Guessing
Trupti Gavit
Founder, KaryoWorks
The ROI problem
Ask any team "is this AI tool worth it?" and you will get one of two answers:
- 1."Absolutely!" (no data)
- 2."I think so..." (no data)
AI ROI is treated like a feeling rather than a calculation. But it is math — if you have the right inputs and subtract the costs that most ROI models ignore.
The scale of the measurement gap is significant. An IBM CEO study found that only 25% of AI initiatives delivered expected ROI over the past three years, and only 16% scaled enterprise-wide. IBM's Think Circle research reports that just 29% of executives can measure AI ROI confidently — while 79% see productivity gains they cannot translate into financial impact.
Why most AI ROI calculations fail
Productivity without conversion
IBM's research notes that 79% of organizations observe productivity gains from AI — but converting those gains into financial impact requires explicit workflow design. Time saved on drafting does not become ROI until that time is redeployed to higher-value work, headcount is avoided, or output volume increases.
Pilot ROI does not scale
IBM's AI research found that as AI pilots scale to production, returns moderate to roughly 7% ROI on average — below the typical 10% cost-of-capital hurdle rate. The top decile of organizations achieves approximately 18% ROI. The difference is not better models — it is better measurement, workflow redesign, and organizational integration.
Enterprise impact vs. tool impact
McKinsey's State of AI 2025 reports that 88% of organizations use AI, but only 39% report EBIT impact at the enterprise level. Individual tool ROI can be positive while enterprise ROI remains invisible — because tools are deployed without connecting to business outcomes.
The formula
Monthly AI ROI = Total Monthly Value Generated - Monthly Cost
Where:
Total Monthly Value = Time Savings Value + Quality Improvement Value + Error Reduction Value - Rework Cost
The subtraction matters. Most ROI models ignore it.
Productivity benchmarks from published research
Before estimating time savings from gut feel, anchor your assumptions. The Alice Labs AI Automation ROI Benchmark 2026 synthesizes peer-reviewed and enterprise case data:
| Use case | Documented gain | Source type |
|---|---|---|
| Customer support | 15% productivity gain | Peer-reviewed (5,179 agents) |
| Professional writing | 40% faster completion | Controlled study (Noy & Zhang) |
| Coding (controlled task) | 55.8% faster completion | GitHub Copilot experiment |
| Knowledge work (suitable tasks) | 12.2% more tasks, 25.1% faster | HBS/BCG jagged-frontier study |
| Copilot-style deployments | 1.9–4.0 hours saved per week | Cross-case enterprise data |
How to use these benchmarks
These are ceiling references, not guarantees. The HBS/BCG study also found 19 percentage points worse correctness outside the AI "frontier" — meaning gains on suitable tasks coexist with losses on unsuitable ones.
For ROI modeling, use the lower bound of each range for conservative estimates and apply benchmarks only to the specific task type — not the whole job. Adjust for your team's adoption rate (enterprise average: 34% per Zylo data).
Time savings value
This is usually the largest component — and the easiest to overestimate.
Formula: Users x Hours saved per week x Hourly rate x 4.33 weeks/month
Conservative benchmark: Use 1.9 hours per week (lower bound from Alice Labs) multiplied by your adoption rate.
Example at full adoption: 5 content writers save 1.9 hours/week each. Loaded hourly rate: $45.
Time Savings Value = 5 x 1.9 x $45 x 4.33 = $1,852/month
Compare this to an estimate using 3 hours/week ($2,924/month). The benchmark-backed number is 37% lower — and more defensible in a finance review.
Quality improvement value
Quality gains are real but harder to quantify. Alice Labs documents 40% faster completion on bounded professional writing with higher rated quality on average. But "faster" is not automatically "better."
Assign a dollar value only where you can quantify it. If quality "feels better," estimate $0 until measured.
Error reduction value
Some AI tools prevent mistakes with real costs: AI code review catching bugs, AI validation catching data entry errors, AI grammar tools preventing client-facing errors.
Formula: Errors caught per month x Average cost per error
Example: AI code review catches 4 bugs/month that would each take 3 hours to fix in production. Developer rate: $75/hour.
Error Reduction Value = 4 x 3 x $75 = $900/month
Only include errors you can document.
Net productivity: Subtract rework and verification time
This is the adjustment most ROI calculations miss.
AI output requires verification. Drafts need editing. Code needs review. Data needs checking. The time AI saves on generation is partially consumed by the time humans spend confirming accuracy.
Why net productivity matters
Alice Labs cites the HBS/BCG jagged-frontier finding: AI produces 19 percentage points worse correctness outside its capability frontier. If your team uses AI on tasks beyond that frontier, verification time increases — potentially erasing time savings.
How to calculate rework cost
Rework Cost = Verification hours per week x Users x Hourly rate x 4.33
| AI output quality | Verification time as % of generation time saved |
|---|---|
| Minor edits only | 10–20% of time saved |
| Regular major revisions | 30–50% of time saved |
| Frequent rejection/regeneration | 50–80% of time saved |
Example: AI saves 1.9 hours/week per writer, but verification consumes 0.5 hours/week.
Net time saved = 1.9 - 0.5 = 1.4 hours/week (26% rework tax)
Adjusted monthly value = 5 x 1.4 x $45 x 4.33 = $1,364/month (vs. $1,852 gross)
See Your AI Is Lying to You for the five most common hallucination patterns that drive rework.
For ROI purposes, treat verification time as a real cost:
- Low-risk tasks (internal brainstorming, rough drafts): 10% rework tax
- Medium-risk tasks (client emails, reports with data): 25% rework tax
- High-risk tasks (published content, legal, financial): 40–50% rework tax
Putting it together: Benchmark-backed example
Tool: AI writing assistant for a 5-person content team. Monthly cost: $300
| Component | Gross estimate | Rework adjustment | Net value |
|---|---|---|---|
| Time savings (1.9 hrs/wk, Alice Labs lower bound) | $1,852 | -26% verification tax | $1,364 |
| Quality improvement (revision rate -20%) | $180 | — | $180 |
| Error reduction | $0 | — | $0 |
| Total monthly value | $2,032 | $1,544 | |
| Monthly cost | $300 | ||
| Monthly ROI | $1,244 | ||
| ROI percentage | 415% |
Even with conservative benchmarks and rework subtracted, a well-fit tool shows strong ROI. Compare to a poorly-fit tool:
| Component | Value |
|---|---|
| Time savings (5 users, 0.3 hrs/wk actual, 40% adoption) | $117 |
| Rework (high rejection rate) | -$58 |
| Quality improvement | $0 |
| Net monthly value | $59 |
| Monthly cost | $300 |
| Monthly ROI | -$241 (negative) |
Same formula. Different inputs. The math tells you what to cut.
When ROI is negative
If monthly cost exceeds net monthly value, check before cancelling:
- 1.Is adoption the problem? At 34% average utilization (Zylo), most tools underperform on adoption before they underperform on capability.
- 2.Is the use case wrong? A powerful tool applied to the wrong task will show usage without value.
- 3.Is it a timing issue? Some tools need 3–6 months for proficiency. But 44% of AI licenses are abandoned within 90 days.
- 4.Is rework eating the savings? Recalculate with verification time included.
Rule: Negative net ROI after 6 months with adoption above 40% = cut the tool.
The sensitivity check
Always stress-test your assumptions. Optimistic ROI models are how 75% of AI initiatives fail to deliver expected returns (IBM).
Run three scenarios:
| Scenario | Time savings assumption | Rework tax | When to use |
|---|---|---|---|
| Conservative | Lower bound benchmark x adoption rate | 30–50% | Default for finance review |
| Expected | Mid-range benchmark x adoption rate | 20–30% | Internal planning |
| Optimistic | Upper bound x 100% adoption | 10–15% | Vendor business case only |
Decision rules:
- ROI positive in conservative scenario = KEEP with confidence
- ROI positive only in expected scenario = OPTIMIZE (improve adoption or reduce rework)
- ROI positive only in optimistic scenario = CUT or REPLACE
- ROI negative in all scenarios = CUT immediately
Context for what "good" looks like at scale: IBM's research puts average scaled AI ROI at 7%, with top-decile organizations at 18%. A tool-level ROI of 200–400% on a specific workflow is plausible. Enterprise-wide AI ROI claims above 18% require exceptional evidence.
Start calculating
Most businesses have never run this calculation for any AI tool. Doing it once for your three most expensive subscriptions — with benchmark-backed inputs, rework subtracted, and sensitivity tested — will tell you more than six months of "I think it's helping."
The AI Automation Audit System includes a pre-built ROI Calculator spreadsheet that automates these formulas. Or start with the free AI Spend Calculator to establish your cost baseline.
For warning signs that you need this calculation now, see 5 Signs You're Wasting Money on AI.
Sources
Trupti Gavit
Founder, KaryoWorks
AI practitioner building evaluation and decision systems for businesses managing AI investments.
Related articles
The AI Subscription Trap: How to Know If Your AI Tools Are Worth It
51% of SaaS licenses go unused. 44% of AI licenses are abandoned within 90 days. A five-dimension framework for evaluating every AI tool — with sourced benchmarks and a worked example.
Read5 Signs You're Wasting Money on AI
88% of organizations use AI but only 39% report financial impact. Five warning signs your AI spend is underperforming — each backed by industry data — with specific action thresholds.
ReadHow to Audit Your AI Stack in One Afternoon
A step-by-step preview of the methodology behind the AI Automation Audit System. Enough to get started, detailed enough to be useful.
ReadGet practical AI frameworks by email
Evaluation systems, spend analysis, and decision tools — one email when we publish. No spam.