Methodology(Updated )· 12 min read

How to Audit Your AI Stack in One Afternoon

TG

Trupti Gavit

Founder, KaryoWorks

Why audit your AI stack now

Most businesses did not plan their AI stack — they accumulated it. One team adopted a writing assistant. Sales added a CRM copilot. Someone expensed ChatGPT Plus. Finance approved an enterprise contract nobody tracks centrally. Six months later, you have overlapping tools, invisible spend, and no shared answer to a basic question: Is any of this working?

The scale of the problem is now measurable. According to Torii's 2026 SaaS Benchmark, the average organization runs 831 SaaS applications, and 61.3% operate outside IT oversight — what the industry calls shadow IT. Of the 50 most common shadow applications, 26 are AI tools. Meanwhile, VendorBenchmark's SaaS spend research finds that 34% of SaaS budget is wasted on unused licenses, redundant tools, and poor procurement.

The difference between "reviewing your AI tools" and auditing your AI stack is methodology. A review is subjective — you look at each tool and decide if it feels worth keeping. An audit is systematic: consistent criteria, collected data, documented decisions. Reviews are influenced by who adopted first, who speaks loudest, and who is most attached to a subscription. Audits produce the same result regardless of office politics — which is exactly why finance and leadership need one before the next renewal cycle.

This guide walks through a five-phase audit you can complete in one afternoon. It is a preview of the methodology behind the AI Automation Audit System, but complete enough to run on your own this week.

The five-phase audit framework

PhaseTimePurpose
1. Inventory~30 minFind every tool, including shadow AI
2. Baseline30–60 minDocument pre-AI performance
3. Evaluate1–2 hoursScore each tool on five criteria
4. Decide~30 minAssign keep / optimize / replace / cut
5. Plan~30 minBuild a 30/60/90-day action plan

The phases build on each other. Skip inventory and you miss hidden spend. Skip baseline and your evaluation becomes opinion. Skip scoring and your decisions revert to politics.

Phase 1: Inventory — including shadow AI (30 minutes)

Why this phase matters: You cannot optimize what you cannot see. Redress Compliance's 2026 Shadow AI report estimates that shadow AI accounts for 4–9% of total software spend — and can run 2–3× the formal AI budget when personal subscriptions, expensed tools, and team-level trials are included. Torii reports that 61.3% of SaaS apps are shadow IT, meaning they were purchased without central visibility. If your inventory only covers IT-approved contracts, you are auditing a fraction of reality.

Open a spreadsheet. For every AI tool your business uses — approved or not — capture:

  • Tool name and vendor
  • Monthly cost (subscription, per-seat, usage-based, or expensed)
  • Licensed users vs. active users (last 30 days)
  • Primary use case (one sentence)
  • Who owns the subscription and who pays the invoice
  • Procurement path: IT-approved, departmental, or individual/expensed

How to surface shadow AI in one afternoon:

  • Pull the last 90 days of corporate card and expense reports; filter for AI-related vendors
  • Ask each department head: "What AI tools does your team use that IT doesn't manage?"
  • Check SSO logs and browser extension reports if available
  • Review Slack/Teams channels for shared links to AI products

Shadow tools are often the most revealing findings — not because they are wrong, but because they represent real workflow demand that central planning missed. Your job in Phase 1 is visibility, not judgment.

Phase 2: Baseline — the step almost everyone skips (30–60 minutes)

Why this phase matters: Without a before-state, you cannot prove after-state improvement — and that is why most organizations struggle to justify AI spend. IBM's AI ROI research found that only 29% of organizations can measure AI ROI confidently. The rest are guessing, which makes budget conversations defensive instead of data-driven.

For each tool's primary use case, document what the process looked like before AI:

  • How long did the task take (minutes or hours per unit)?
  • What was the error rate or revision rate?
  • What was the output volume per week or month?
  • What did it cost in labor, rework, or outsourced fees?

If you lack historical data, estimate conservatively and label it as an estimate. A rough baseline beats no baseline — and you can refine it during Phase 3.

Why teams skip this: Baselines require talking to the people doing the work, not the people who bought the license. That friction is the point. An audit that only reviews invoices tells you what you spend. An audit with baselines tells you whether spending changed outcomes.

Without baselines, Phase 3 produces scores based on enthusiasm. With baselines, it produces scores based on evidence.

Phase 3: Evaluate — five criteria with utilization benchmarks (1–2 hours)

Why scoring matters: Subjective review collapses under disagreement. A weighted score gives every stakeholder the same language. It also connects to industry utilization data: Lynton's AI spend audit framework (drawing on Zylo benchmarks) reports 34% average AI tool utilization and 44% of AI licenses abandoned within 90 days. Those numbers explain why many stacks feel productive in demos but hollow in practice — low adoption masquerades as tool failure, and high adoption without impact masquerades as success.

For each tool, score 1–5 on:

  1. 1.Task-AI fit — Is this the right kind of task for AI? (Repetitive, pattern-based, language-heavy tasks score higher; judgment-heavy or relationship-critical tasks score lower.)
  2. 2.Adoption — What percentage of intended users actually use it weekly?
  3. 3.Impact — Is there measurable improvement vs. your Phase 2 baseline?
  4. 4.Alternative cost — Is this tool the best option for the job, including free or bundled alternatives?
  5. 5.Risk — What is the dependency risk if this vendor changes pricing, terms, or availability?

Calculate a weighted average per tool. Use these thresholds, aligned with utilization benchmarks from SaaS management practice:

Score / signalRecommended action
Weighted score above 3.5 and adoption above 60%Strong KEEP candidate
Score 2.5–3.5 or adoption 30–60%OPTIMIZE — fix onboarding, use case, or ownership
Score below 2.5 and adoption below 30%CUT or REPLACE — licenses are likely shelfware
High score but low adoptionProcess problem, not tool problem — assign an owner before cutting

Tools scoring below 3.0 are action candidates. Tools with adoption below 30% after 90 days align with the 44% license abandonment rate — default to skepticism unless impact scores are exceptionally high.

Phase 4: Decide — validate with cost benchmarks (30 minutes)

Why decisions need external anchors: Internal scores tell you relative value. Industry benchmarks tell you whether your total spend is reasonable. Without both, teams cut the wrong tools or keep expensive duplicates.

Based on Phase 3 scores and your judgment, assign each tool one of four labels:

  • KEEP — Score above 3.5, adoption above 60%, clear measured value
  • OPTIMIZE — Score 2.5–3.5, real need exists but implementation is weak
  • REPLACE — Low score but the underlying workflow need is real; evaluate alternatives
  • CUT — Low score, low adoption, and the need is questionable or duplicated elsewhere

Use these benchmarks to pressure-test your decisions:

BenchmarkSourceWhat it tells you
34% of SaaS budget wastedVendorBenchmarkIf you cannot identify ~⅓ of spend as low-value, you have not looked hard enough
23% of seats inactiveVendorBenchmarkAny tool with >20% inactive seats is an OPTIMIZE or CUT candidate
28% of spend is shadow ITVendorBenchmarkUncatalogued tools should not automatically become KEEP
Shadow buyers pay 35–60% moreVendorBenchmarkREPLACE decisions should include central procurement repricing

If your KEEP list includes five tools doing substantially the same job, the benchmark data says you are probably wrong — redundancy is one of the largest drivers of the 34% waste figure.

Phase 5: Plan — decisions plus governance (30 minutes)

Why governance belongs in the plan: Cutting tools without guardrails invites shadow AI to return within a quarter. Precisely's 2026 State of Data Integrity and AI Readiness found that organizations with formal data governance report 71% data trust versus 50% without governance — a 21-point gap that directly affects whether AI tools produce reliable outputs. An audit plan that only cancels subscriptions, without approval paths and data rules, treats symptoms.

Document your Phase 4 decisions and build a 30/60/90-day plan:

  • 30 days: Cancel CUT tools; reassign or downgrade inactive seats; begin REPLACE evaluations for high-need gaps; publish a simple AI tool request process
  • 60 days: Implement OPTIMIZE changes (training, workflow integration, ownership); trial replacement candidates; consolidate shadow tools into approved alternatives where possible
  • 90 days: Re-score OPTIMIZE tools; measure adoption and impact against Phase 2 baselines; report savings and governance metrics to leadership

Minimum governance to add in Phase 5:

  • A single owner for AI stack visibility (often IT + finance shared)
  • An approved-tool list with a lightweight exception process
  • Data handling rules for tools that process customer or proprietary information
  • A quarterly re-audit cadence (shorter than annual — AI tools churn fast)

What this afternoon audit produces

At the end of one focused afternoon, you should have:

  • A complete inventory — including shadow AI that never appeared on IT's radar
  • A cost baseline — total spend, per-user spend, and shadow vs. approved split
  • A performance baseline — pre-AI metrics for each major use case
  • A scored evaluation — every tool rated on the same five criteria
  • Documented decisions — keep, optimize, replace, or cut for each tool
  • A 90-day action plan — with governance to prevent re-sprawl

Expected savings range: Organizations with active SaaS management typically waste 8–15% of software budget, according to VendorBenchmark. Organizations without centralized visibility waste 45–60%. An afternoon audit will not capture all of that gap immediately — but most teams identify 10–20% quick-win reduction in the first cycle through inactive seat recovery, duplicate tool elimination, and shadow-to-approved consolidation. Deeper savings come from OPTIMIZE and REPLACE work in the 60–90 day window.

That is more structured output than most organizations produce after months of "thinking about AI strategy."

Common mistakes that undermine the audit

Even teams that commit to an audit often undermine their own results. Watch for these patterns:

Auditing only IT-approved tools. If Phase 1 excludes expense reports and departmental purchases, you will "save" 10% while missing the 28% of spend tied to shadow IT. Finance should co-own Phase 1, not just review the output.

Skipping baselines because data is imperfect. Imperfect baselines still beat none. IBM links low ROI confidence directly to weak measurement foundations — and organizations without governance report 50% data trust vs. 71% with governance. Estimate, label, refine.

Letting the loudest advocate veto CUT decisions. Scoring exists to override politics. When a tool scores below 2.5 with adoption under 30%, the 44% 90-day abandonment rate is your external justification — not personal preference.

Treating the audit as a one-time event. AI stacks change monthly. The Phase 5 quarterly re-audit cadence is not optional if you want to stay inside the 8–15% waste band instead of drifting back toward 45–60%.

Going deeper

This article gives you the process. The AI Automation Audit System gives you the tools: pre-built spreadsheets with scoring formulas, ROI calculators, decision matrices, report templates, and a methodology guide that goes deeper than a single blog post. For the strategy layer — deciding what to adopt next, not just what to cut — see Stop Collecting AI Tools and the AI Decision System.

Start with Phase 1 this week. The shadow AI line items alone usually pay for the afternoon.

Sources

TG

Trupti Gavit

Founder, KaryoWorks

AI practitioner building evaluation and decision systems for businesses managing AI investments.

Get practical AI frameworks by email

Evaluation systems, spend analysis, and decision tools — one email when we publish. No spam.